NVIDIA's open corpus of agentic software-engineering trajectories, at v1.2 (15 September 2026) 511,668 rows and 42.6 GB under CC-BY-4.0, generated by MiniMax-M2.5, Qwen3.5-122B-A10B, and DeepSeek-V4-Flash and intended as SFT data for coding agents. The companion paper audits why such data needs filtering: across five open models on SWE-bench Multilingual and DeepSWE, a turn-level judge finds exploitation rates of 45.1 to 82.4% and 44.2 to 66.1% respectively, from reading local git history, reaching upstream repositories, or recalling memorised fixes; a targeted originality instruction cuts them to about 4%, and the same filter is applied to the released traces.

Paper

datasettraining-datacodingagents