Sudeep Pillai PRO
spillai
AI & ML interests
Self-supervised learning, Few-shot learning, Computer Vision, Robotics
Recent Activity
reacted to nwaughachukwuma's post with 🔥 1 day ago
# VLM Run Gateway: Run GLM-OCR, DeepSeek-OCR-2, and Dots.mocr with an OpenAI Compatible API
Open-weight OCR VLMs have advanced significantly over the past year, yet most teams still rely on frontier VLMs for document parsing because researching, evaluating, and deploying the right models remains challenging.
So we built VLM Run Gateway: one OpenAI-compatible endpoint for open-weight OCR and VLM models.
If you’re using frontier VLMs primarily for OCR/document parsing, open-weight OCR models can be dramatically cheaper and often very accurate. With a one-line change, you can switch between open-weight OCR VLMs (DeepSeek OCR 2, GLM-OCR, dots.mocr, Paddle OCR VL, PP-OCRv6, etc.) and process 100K+ pages for under $60.
Try it out quickly via OpenAI SDK:
```
client = OpenAI(
base_url="https://gateway.vlm.run/v1/openai",
api_key="<VLMRUN_API_KEY>",
)
response = client.chat.completions.create(
model="rednote-hilab/dots.mocr",
messages=[{
"role": "user",
"content": [{
"type": "document_url",
"document_url": {"url": "https://.../invoice.pdf"},
}],
}],
extra_body={"document_dpi": 72},
)
```
or via our CLI:
```
pip install vlmrun
vlmrun gw models
vlmrun config set --api-key 'vlmrun' # anon-user, rate-limited
vlmrun gw chat <doc>.pdf -m zai-org/glm-ocr
vlmrun gw chat <doc>.pdf -m zai-org/glm-ocr --json-mode
vlmrun gw chat <doc>.pdf -m deepseek-ai/deepseek-ocr-2
vlmrun gw chat <doc>.pdf -m rednote-hilab/dots.mocr
vlmrun gw chat <doc>.pdf -m paddleocr/pp-ocrv6
```
Docs: https://docs.vlm.run/gateway
Catalog: https://docs.vlm.run/gateway/models
MCP: https://docs.vlm.run/gateway/mcp-server
Colab Quickstart: https://colab.research.google.com/drive/1RkuVIyuc5Po-UlcSlFyJCam5tjCm9IHM?usp=sharing
Read the full post here: https://huggingface.co/blog/vlm-run/intro-to-vlmrun-gateway new activity 16 days ago
huggingface/InferenceSupport:dots-studio/dots.mocr reacted to nwaughachukwuma's post with 🤗 17 days ago
Can a text-only model + a vision toolkit (mm-ctx) match a native vision model?
We benchmarked 4 setups on 23 multimodal tasks (image, video, audio, PDF):
• glm-5.2 (text-only) + mm-ctx: 88.4
• gemini-3.5-flash (vision): 83
• deepseek-v4-pro (text-only) + mm-ctx: 79.4
• qwen3.6-35b-a3b (vision): 44.3
The best text-only setup `glm-5.2 + mm` outperformed gemini-3.5-flash, the top vision model, by 5.4 points (6.5%). It was also:
• 1.5x faster (100s vs 150s mean per task)
• the only setup with zero timeouts (46/46 completed; gemini timed out 4x on bulk-image and long-video tasks)
• the only setup stable across runs (88.5 / 88.4)
• top on video (100.0), image (91.7), and PDF (90.0) tasks
The trade-offs: the toolkit consumed 3.3x more tokens (4.25M vs 1.28M), and lost on audio (85.6 vs 71.3).
On completed tasks alone the two are nearly identical (91.0 vs 88.4): the toolkit's edge is efficient extraction that keeps long media tasks inside the time budget.
Full report: https://huggingface.co/blog/vlm-run/text-only-models-with-mm