We are featured on Product Hunt today!Support Us & Vote on Product Hunt ↗
ModelVaultModel Finder
All Models Index11k+Cloud API ModelsAPIsLocal Models (Ollama)Free
Find a modelNewPlaygroundWebGPUTelemetry ⚡
AI Cost CalculatorCalculatorVRAM Hardware EstimatorGPUToken CounterTokensContext CapacityContextFind a modelFinder
SavedCompare
Compare
Models/Yolos Small
HustvlOpen Weights✓ Verified Spec & Code

Yolos Small

Open Weights • Released 2022-04-26 • Last Verified 2026-08-06

Try PlaygroundDocs

Yolos Small is a high-performance multimodal vision-language model developed by Hustvl. Integrating advanced visual encoder networks with deep language models, Yolos Small excels at visual document understanding (DocVQA), chart and diagram parsing, high-resolution optical character recognition (OCR), and spatial reasoning.

Context Window128k
LicenseApache-2.0
Deployment
Local Model
API AvailableYes (REST/SDK)
💡

Plain English Summary (What is this model & who is it for?)

Think of Yolos Small as an AI with eyes. You can upload photos, receipts, financial charts, or scanned documents, and ask it to read text, analyze visual contents, or answer questions about what it sees.

💡 Real-World Use Cases & Practical Examples

📄 Document & Receipt Parsing: Extract total amounts, dates, and line items from invoices and receipts.
📊 Chart & Graph Analysis: Upload financial reports to extract trends and summary insights.
🔍 Visual Inspection: Identify products, serial numbers, or visual damage in uploaded photos.
🖼️ Image Captioning: Generate detailed alt-text and descriptions for accessibility and SEO.

🚀 How to Run & Use This Model (Step-by-Step Guide)

Simple setup instructions for everyday users and developers.

1

Download a One-Click App (No Coding Required)

Download a free local AI launcher like LM Studio (lmstudio.ai) or Ollama (ollama.com) on your Mac, Windows, or Linux PC.

2

Load the Model

In LM Studio, search for "Yolos Small". In Ollama, open your terminal and run "ollama run hustvl-yolos-small".

3

Start Chatting or Generating

Type your text instructions or upload files into the app. The AI runs 100% privately on your hardware without internet requirement!

4

Developer API Integration

Developers can integrate Yolos Small directly via Python (using Hugging Face transformers/diffusers) or connect via local OpenAI-compatible REST server (http://localhost:11434).

Benchmark Performance

DocVQA (Document QA)87.9
MMBench (Multimodal)82.4
MathVista (Visual Math)73.8
ChartQA Score84.1

Hardware Requirements for Local Running

Requires dedicated GPU with 16GB+ VRAM recommended for fast local inference.

Strengths

  • •High-Resolution Document VQA & OCR
  • •Multi-Chart & Diagram Structural Parsing
  • •Spatial Object Detection & Annotation
  • •Seamless Text-Image Input Fusion

Limitations & Weaknesses

  • •Requires local GPU hardware for self-hosting
Integration Code (vision-language)
from transformers import AutoProcessor, AutoModelForVision2Seq
from PIL import Image
import torch

model_id = "hustvl/yolos-small"
model = AutoModelForVision2Seq.from_pretrained(model_id, torch_dtype=torch.float16, device_map="auto")
processor = AutoProcessor.from_pretrained(model_id)

image = Image.open("sample.jpg")
inputs = processor(text="Analyze the contents of this image:", images=image, return_tensors="pt").to("cuda")

outputs = model.generate(**inputs, max_new_tokens=150)
print(processor.batch_decode(outputs, skip_special_tokens=True)[0])

Pricing Overview

Free Open Weights / Self-Hosted

Prices subject to provider tiers and volume discounts. Check documentation for current token rates.

Model Tags

#transformers#pytorch#safetensors#yolos#object-detection#vision

Did this model work for you?

Your feedback helps others find the right model.

ModelVault

Find a suitable AI model for your task, budget and hardware, with sources and practical setup guidance.

Curated recommendations · Source-backed specifications
ProductAll AI Models IndexComparison MatrixLocal Models (Ollama)Cloud API ModelsFind a modelLive API Telemetry ⚡WebGPU AI Playground 🎮AI Cost Calculator
ResourcesReasoning ModelsCoding AgentsVision-LanguageEmbeddings & RAG
Management & LegalAdmin Console ↗Privacy PolicyTerms of ServiceDisclaimerAbout ModelVaultContact Us
© 2026 ModelVault AI Directory. Review model sources before deployment.
Made with❤️by Gaurav Kushwaha
Built for high-performance AI workflows