TIPSv2 (CVPR'26) and TIPS (ICLR'25)
-
Updated
Aug 31, 2026 - Jupyter Notebook
TIPSv2 (CVPR'26) and TIPS (ICLR'25)
OpenVision (ICCV 2025), OpenVision 2 (CVPR 2026), and OpenVision 3
timm, evolved
A Gradio-based demonstration for the AllenAI Molmo2-8B multimodal model, enabling image QA, multi-image pointing, video QA, and temporal tracking. Users upload images or videos, provide natural language prompts.
[NeurIPS 2026] Implementation of the paper "Vision Correlators: Correlation-Driven Visual Understanding with Hypergraphs".
This demonstrates the process of adapting a large scale pretrained model, MetaCLIP 2, for fine tuning a specific downstream task: image classification.
Coral-Health is an image classification vision-language encoder model fine-tuned from google/siglip2-base-patch16-224 for a single-label classification task. It is designed to classify coral reef images into two health conditions using the SiglipForImageClassification architecture.
Flood-Image-Detection is a vision-language encoder model fine-tuned from google/siglip2-base-patch16-512 for binary image classification. It is trained to detect whether an image contains a flooded scene or non-flooded environment. The model uses the SiglipForImageClassification architecture.
Multilabel-GeoSceneNet is a vision-language encoder model fine-tuned from google/siglip2-base-patch16-224 for multi-label image classification. It is designed to recognize and label multiple geographic or environmental elements in a single image using the SiglipForImageClassification architecture.
Self-supervised monocular depth estimation using a Vision Transformers [Encoder, multi-scale CNN decoder, pose estimation, and differentiable view reconstruction].
Multilabel-Portrait-SigLIP2 is a vision-language model fine-tuned from google/siglip2-base-patch16-224 using the SiglipForImageClassification architecture. It classifies portrait-style images into one of the following visual portrait categories:
shoe-type-detection is a vision-language encoder model fine-tuned from google/siglip2-base-patch16-512 for multi-class image classification. It is trained to detect different types of shoes such as Ballet Flats, Boat Shoes, Brogues, Clogs, and Sneakers. The model uses the SiglipForImageClassification architecture.
Benchmarking SSM Based Image Encoder Performance in Solar Irradiance Prediction using All-Sky Images
Fashion-Product-Usage is a vision-language model fine-tuned from google/siglip2-base-patch16-224 using the SiglipForImageClassification architecture. It classifies fashion product images based on their intended usage context.
PussyCat-vs-Doggie-SigLIP2 is an image classification vision-language encoder model fine-tuned from google/siglip2-base-patch16-224 for a single-label classification task. It is designed to classify images as either a cat or a dog using the SiglipForImageClassification architecture.
Leverage SigLIP 2's capabilities using LitServe.
Research project for exploring Kazakh Vision Language Models.
💠 Ambiente local e aberto, com ferramentas para observar, experimentar, inspecionar e documentar representações visuais de modelos. 💠 Primeira bancada: 🔹 Bancada Visual Qwen3.5-9B: ferramenta local para preparar imagens e sequências e explorar representações da torre visual do Qwen3.5-9B, com os pesos congelados.
Cloudflare Clef 方向的多模态决策模型框架:指针决策头(选项打分替代词表生成)+ 原生视觉编码器 + 置信度校准门控。零依赖纯 Python 参考实现
To associate your repository with the vision-encoder topic, visit your repo's landing page and select "manage topics."