{"rewrite":{"id":"r_dc253705c0d5b7b79139e3f3","clusterId":"c_499db7388ac8735d293ade67","slug":"zyphra-releases-zamba2-vl-vision-language-models-with-hybrid-ssm-transformer-architecture","model":"deepseek-v4-flash","headline":"Zyphra Releases Zamba2-VL Vision-Language Models With Hybrid SSM-Transformer Architecture","summary":"Zyphra released the Zamba2-VL family of vision-language models on June 11, 2026. The models use a hybrid SSM-Transformer architecture that combines Transformer with Mamba2, aiming for faster image recognition at quality comparable to similarly scaled Transformer models. Three variants are available: 1.2B, 2.7B, and 7B parameters, all open under Apache License 2.0.","whyItMatters":"Zamba2-VL offers a practical alternative to pure Transformer VLMs by trading architectural complexity for speed without sacrificing benchmark scores, and releasing all sizes openly lowers the barrier for developers who need fast image recognition on limited hardware.","webCardHtml":"\u003cp\u003eZyphra announced the Zamba2-VL family of vision-language models on June 11, built on a hybrid architecture the company calls SSM-Transformer. The design combines the dominant Transformer architecture with Mamba2, a state-space model introduced in 2024. Zyphra claims the hybrid achieves faster image recognition processing than Transformer-only models of similar scale while maintaining equivalent quality on benchmarks.\u003c/p\u003e\u003cp\u003eThree model sizes are released: Zamba2-VL-1.2B (2 billion parameters), Zamba2-VL-2.7B (2.7 billion), and Zamba2-VL-7B (7 billion). All three are open models distributed under the Apache License 2.0 and available for download via Hugging Face. The company published a graph showing time to first token against average benchmark score, positioning the Zamba2-VL series as competitive on both speed and accuracy relative to peers.\u003c/p\u003e\u003cul\u003e\u003cli\u003e\u003cstrong\u003eZamba2-VL-1.2B\u003c/strong\u003e: 2-billion-parameter variant\u003c/li\u003e\u003cli\u003e\u003cstrong\u003eZamba2-VL-2.7B\u003c/strong\u003e: 2.7-billion-parameter variant\u003c/li\u003e\u003cli\u003e\u003cstrong\u003eZamba2-VL-7B\u003c/strong\u003e: 7-billion-parameter variant\u003c/li\u003e\u003c/ul\u003e","blueskyPost":"Zyphra released Zamba2-VL, a family of vision-language models using a hybrid SSM-Transformer architecture. Three sizes from 1.2B to 7B parameters, all open under Apache 2.0. Faster image recognition than similarly scaled Transformer models, per the company.","twitterPost":"Zyphra released Zamba2-VL, a family of vision-language models using a hybrid SSM-Transformer architecture. Three sizes from 1.2B to 7B parameters, all open under Apache 2.0. Faster image recognition than similarly scaled Transformer models, per the company.","threadsPost":null,"newsletterBlurb":"Zyphra released the Zamba2-VL family of vision-language models on June 11. The models use a hybrid SSM-Transformer architecture that combines Transformer with Mamba2, aiming for faster image recognition at quality comparable to similarly scaled Transformer models. Three variants are available: 1.2B, 2.7B, and 7B parameters, all open under Apache License 2.0.","attributionJson":"[{\"source\":\"GIGAZINE\",\"url\":\"https://gigazine.net/news/20260611-zamba2-vl-zyphra/\",\"title\":\"High-Speed and High-Precision Vision-Language Model \\\"Zamba2-VL\\\" Appears, Developed with Architecture Faster Than Transformer\"}]","lintFlagsJson":null,"lintHits":0,"costUsd":0,"inputTokens":4296,"outputTokens":794,"status":"published","repairAttempts":0,"nextRepairAt":null,"factsAttemptedAt":1781171620,"createdAt":"2026-06-11T09:41:58.000Z","publishedAt":"2026-06-11T09:45:33.000Z","updatedAt":"2026-06-11T09:45:33.000Z"},"cluster":{"id":"c_499db7388ac8735d293ade67","canonicalTitle":"高速かつ高精度な視覚言語モデル「Zamba2-VL」が登場、Transformerより高速なアーキテクチャで開発","representativeArticleId":"a_47e80621d6ed39e695829add","sourceCount":1,"writtenSourceCount":1,"writeAttempts":0,"isSolo":true,"entitiesJson":"{\"anime_titles\":[],\"manga_titles\":[],\"work_titles\":[],\"studios\":[],\"people\":[],\"type\":\"news\",\"domain\":\"other\",\"is_roundup\":false}","contentType":"news","status":"published","firstSeenAt":"2026-06-11T08:49:00.000Z","lastSeenAt":"2026-06-11T08:49:00.000Z","updatedAt":"2026-06-11T09:45:33.000Z"},"attribution":[{"source":"GIGAZINE","url":"https://gigazine.net/news/20260611-zamba2-vl-zyphra/","title":"高速かつ高精度な視覚言語モデル「Zamba2-VL」が登場、Transformerより高速なアーキテクチャで開発"}],"entities":{"anime_titles":[],"manga_titles":[],"work_titles":[],"studios":[],"people":[],"type":"news","domain":"other","is_roundup":false},"keyFacts":["Zyphra released the Zamba2-VL family of vision-language models on June 11, 2026.","The models use a hybrid SSM-Transformer architecture that combines Transformer with Mamba2.","Three variants are available: 1.2B, 2.7B, and 7B parameters.","All three models are open under Apache License 2.0 and available on Hugging Face.","Zyphra claims the hybrid architecture achieves faster image recognition than Transformer-only models of similar scale."]}
