Agents-A1
model Your tags
Your notes
"Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent" — InternScience's agentic VLM: a 35B-A3B MoE (hybrid linear/full attention, image-text-to-text, Apache-2.0) trained on trajectories averaging ~45K tokens via a three-stage recipe: full-domain SFT, domain-level teacher training, then multi-teacher multi-domain on-policy distillation with heterogeneity-aware optimization. Card-reported results claim SOTA or best-in-class against GPT-5.5, DeepSeek-V4-Pro, and Kimi-K2.6 on agentic and science benchmarks — GAIA 96.0, BrowseComp 75.5, HLE-with-tools 47.6, IFBench 80.6, SciCode 44.3. Strong first-week traction (194 HF likes). Base-model lineage is not stated on the card.
Model Details
Architecture MOE
Parameters 35B
Active params 3B
License Apache 2.0