Open-SWE Adds Qwen Traces
Code Training Enters the Judgment Era
This week scanned 86 HF orgs · 50 GitHub orgs · 71 blogs · 125 X accounts
Open-SWE-Traces added Qwen3.8 traces on September 26, expanding code Agent training data from results to processes [P0]; Surge AI launched DAYJOB Finance and Healthcare on September 25, making enterprise knowledge work judgment an independent benchmark category [P0]; JEV decision models and reproducible experiments emerged in concentration on September 24, indicating that lightweight judgment layers are becoming Agent Infrastructure [P1]. The strongest data demand signal this week: software engineering Agent traces.
Key Findings
This week's 5 high commercial value findings
NVIDIA's Open-SWE-Traces released a v1.2 update on 2026-09-26, adding traces generated by mini-swe-agent powered by Qwen3.8-27B, with plans to continue adding OpenCode and Claude Code harness data. The dataset currently has 31,393 downloads and 136 likes, is licensed under CC-BY-4.0, and covers code, tool calls, and Agent traces.
On 2026-09-25, Surge AI simultaneously released surgeai/DAYJOB-finance and surgeai/DAYJOB-healthcare. Both are designed around real-world enterprise knowledge work, emphasizing short requests, exploration in complex work environments, and tasks requiring professional judgment; both currently have 0 public downloads but have been categorized as Agent Tool scenarios. In a concurrent blog post, Surge AI defined DAYJOB as a benchmark family focused on real-world work capabilities, values, taste, and implicit judgment.
The paper 《Calibrated Decision Models for Autonomous Penetration-Testing Harnesses》 was released on 2026-09-24, studying how JEV and Laya provide a typed, calibrated decision layer for penetration-testing Agents to reduce false positives and severity inflation. On the same day, Ollaya received 392 upvotes and 108 comments on Hacker News; JevBench, released on September 22, received 147 upvotes and 37 comments. Together released togethercomputer/Tev1-4B-experimental on 2026-09-23, which currently has 1,020 downloads and 16 likes, followed by a 0.8B version on 2026-09-25.
《Context-Continuous Preference Learning for Exoskeleton Personalization》 was released on 2026-09-23, exploring how to learn exoskeleton preferences under different operating conditions using limited user feedback. 《Self-Play Pretraining with Zero Data》 was released on 2026-09-24, proposing that models generate training data most valuable for their own improvement. 《Optimal Sequential Annotations for Off-Policy Evaluation》, published on 2026-09-22, studied how to select the most valuable sequential feedback when expert costs are high and LLM judgment may be biased. 《EDGEGEN》, published on 2026-09-21, explored synthetic boundary-case generation for enterprise Agent states and databases.
Qwen released Qwen/RecreationBench on 2026-09-18, containing 250 application recreation tasks across Ubuntu, macOS, Windows, Android, and the Web. It currently has 4,404 downloads and 12 likes and is used for held-out evaluation in RecreationWorld. Qwen also released Qwen-Image-2.1 on 2026-09-14, which currently has 42,469 downloads and 2,333 likes; its image-editing version was released on 2026-09-20 and has 7,200 downloads.
Demand Signals
Infer training data demands from model releases
Want to discuss this issue?
Auto-generated by AI Dataset Radar · Updated weekly
AI Dataset Radar →