ByteDance Seed
首页模型博客&论文Seed Edge加入我们
首页模型博客&论文Seed Edge加入我们
体验豆包AI 体验中心

2025-05-13

Seed1.5-VL Technical Report

Download PDF
上一篇下一篇

摘要

We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter vision encoder and a Mixture-of-Experts (MoE) LLM of 20B active parameters. Despite its relatively compact architecture, it delivers strong performance across a wide spectrum of public VLM benchmarks and internal evaluation suites, achieving the state-of-the-art performance on 38 out of 60 public benchmarks. Moreover, in agent-centric tasks such as GUI control and gameplay, Seed1.5-VL outperforms leading multimodal systems, including OpenAI CUA and Claude 3.7. Beyond visual and video understanding, it also demonstrates strong reasoning abilities, making it particularly effective for multimodal reasoning challenges such as visual puzzles. We believe these capabilities will empower broader applications across diverse tasks. In this report, we mainly provide a comprehensive review of our experiences in building Seed1.5-VL across model design, data construction, and training at various stages, hoping that this report can inspire further research. Seed1.5-VL is now accessible at this https URL (Volcano Engine Model ID: doubao-1-5-thinking-vision-pro-250428)

作者

Seed Multimodal Team

期刊/会议

arXiv

追求智能上限,创造社会价值
欢迎加入字节跳动 Seed
模型成果
Seed2.1
Seedance 2.5
Seedream 5.0 Pro
SeedRealtime
Seed Audio 1.0
Seed GR-RL
了解更多
论文
Seed Edge
Seed STEM 科学家计划
校园招聘
Copyright© 2026 Bytedance Seed
网站声明联系我们 : seed.feedback@bytedance.com
Seed2.1Seedance 2.5Seedream 5.0 ProSeedRealtimeSeed Audio 1.0Seed GR-RL
论文Seed EdgeSeed STEM 科学家计划校园招聘
追求智能上限,创造社会价值
欢迎加入字节跳动 Seed
Copyright© 2026 Bytedance Seed
网站声明
联系我们 : seed.feedback@bytedance.com
追求智能上限,创造社会价值
联系我们 : seed.feedback@bytedance.com
Copyright© 2026 Bytedance Seed网站声明