Radar Brief Week 30, 2026 · 2026-09-19 — 2026-09-26

Open-SWE Adds Qwen Traces
Code Training Enters the Judgment Era

This week scanned 86 HF orgs · 50 GitHub orgs · 71 blogs · 125 X accounts

0
Valuable Datasets
0
Related Papers
0
Blog Posts
0
Active Repos
One-line Summary

Open-SWE-Traces added Qwen3.8 traces on September 26, expanding code Agent training data from results to processes [P0]; Surge AI launched DAYJOB Finance and Healthcare on September 25, making enterprise knowledge work judgment an independent benchmark category [P0]; JEV decision models and reproducible experiments emerged in concentration on September 24, indicating that lightweight judgment layers are becoming Agent Infrastructure [P1]. The strongest data demand signal this week: software engineering Agent traces.

Key Findings

This week's 5 high commercial value findings

P0 Open-SWE-Traces added Qwen3.8 traces on September 26, expanding code Agent training data from results to processes [P0]

NVIDIA's Open-SWE-Traces released a v1.2 update on 2026-09-26, adding traces generated by mini-swe-agent powered by Qwen3.8-27B, with plans to continue adding OpenCode and Claude Code harness data. The dataset currently has 31,393 downloads and 136 likes, is licensed under CC-BY-4.0, and covers code, tool calls, and Agent traces.

Business significance → The competitive focus has shifted from “whether a model can complete a coding task” to “how a model plans, calls tools, recovers from failures, and completes the task.” High-value opportunities are no longer limited to collecting code answers; they now involve having people with software engineering experience judge whether each modification is reasonable, whether the cause of failure is genuine, and whether the repair path is safe, producing trace data usable for process rewards and error attribution. Knowlyr can prioritize building code Agent trace review, tool-call quality judgment, and long-horizon task debriefing capabilities, enabling people to earn income by contributing professional judgment.
P0 Surge AI launched DAYJOB Finance and Healthcare on September 25, making enterprise knowledge work judgment an independent benchmark category [P0]

On 2026-09-25, Surge AI simultaneously released surgeai/DAYJOB-finance and surgeai/DAYJOB-healthcare. Both are designed around real-world enterprise knowledge work, emphasizing short requests, exploration in complex work environments, and tasks requiring professional judgment; both currently have 0 public downloads but have been categorized as Agent Tool scenarios. In a concurrent blog post, Surge AI defined DAYJOB as a benchmark family focused on real-world work capabilities, values, taste, and implicit judgment.

Business significance → Premium providers are directly productizing “professional human judgment” as Agent evaluation benchmarks, demonstrating that the scarce asset in finance, healthcare, law, sales operations, and other fields is not text volume but standards for judging risk, intent, priorities, and action outcomes. Knowlyr should break enterprise knowledge work into four layers—executable tasks, contextual exploration, professional decision-making, and outcome review—prioritize developing expert judgment communities in finance and healthcare, and build higher-value data products than generic question answering.
P1 JEV decision models and reproducible experiments emerged in concentration on September 24, indicating that lightweight judgment layers are becoming Agent Infrastructure [P1]

The paper 《Calibrated Decision Models for Autonomous Penetration-Testing Harnesses》 was released on 2026-09-24, studying how JEV and Laya provide a typed, calibrated decision layer for penetration-testing Agents to reduce false positives and severity inflation. On the same day, Ollaya received 392 upvotes and 108 comments on Hacker News; JevBench, released on September 22, received 147 upvotes and 37 comments. Together released togethercomputer/Tev1-4B-experimental on 2026-09-23, which currently has 1,020 downloads and 16 likes, followed by a 0.8B version on 2026-09-25.

Business significance → Decisions about “whether to pass, whether to escalate, and whether human intervention is needed” are separating from generative answers and becoming calibrated judgment tasks. Training these decision layers requires large numbers of boundary cases, confidence judgments, risk classifications, and human review outcomes—none of which can be reliably obtained through synthetic answers alone. Knowlyr can enter the market through Agent verifier, risk classification, refusal judgment, and human escalation strategy data, with particular focus on high-risk scenarios such as cybersecurity, healthcare, and enterprise approvals.
P1 Research from September 23 to 24 advanced preference learning and self-generated data in parallel, shifting the data quality bottleneck from “quantity” to “feedback effectiveness” [P1]

《Context-Continuous Preference Learning for Exoskeleton Personalization》 was released on 2026-09-23, exploring how to learn exoskeleton preferences under different operating conditions using limited user feedback. 《Self-Play Pretraining with Zero Data》 was released on 2026-09-24, proposing that models generate training data most valuable for their own improvement. 《Optimal Sequential Annotations for Off-Policy Evaluation》, published on 2026-09-22, studied how to select the most valuable sequential feedback when expert costs are high and LLM judgment may be biased. 《EDGEGEN》, published on 2026-09-21, explored synthetic boundary-case generation for enterprise Agent states and databases.

Business significance → Synthetic data has not eliminated human judgment; it has increased dependence on small amounts of high-quality feedback. People still need to decide which cases are worth reviewing, whether preferences change with context, and whether a model's self-evaluation is trustworthy. The value of data service companies will shift from bulk production to feedback sampling, difficult-case discovery, expert calibration, and synthetic data acceptance. Knowlyr should establish a closed-loop of “model generation—human filtering—expert review—retraining,” rather than selling static datasets independently.
P2 Qwen released RecreationBench on September 18, with open-source models beginning to close Agent capability gaps through verifiable computer-use environments [P2]

Qwen released Qwen/RecreationBench on 2026-09-18, containing 250 application recreation tasks across Ubuntu, macOS, Windows, Android, and the Web. It currently has 4,404 downloads and 12 likes and is used for held-out evaluation in RecreationWorld. Qwen also released Qwen-Image-2.1 on 2026-09-14, which currently has 42,469 downloads and 2,333 likes; its image-editing version was released on 2026-09-20 and has 7,200 downloads.

Business significance → Competition among open-source models is shifting from single-turn text capabilities toward reproducible, real-world operation across devices and interfaces. RecreationBench-style tasks require people to judge whether an interface was understood correctly, whether the operation path matches the objective, and whether the recreated result is usable. These fine-grained human-computer interaction judgments will drive demand for computer-use Agents, GUI operations, and multimodal result acceptance. Knowlyr should prepare for cross-system operation tasks and real user experience evaluation rather than focusing only on image-text question answering.

Demand Signals

Infer training data demands from model releases

Data Type Intensity Trend Related Signals
Software Engineering Agent Traces
Extremely strong ↑ New
Open-SWE-Traces added Qwen3.8-27B traces on September 26, with 31,393 downloads and 136 likes
Enterprise Knowledge Work Judgment
Extremely strong ↑ New
Surge AI launched DAYJOB Finance and Healthcare simultaneously on September 25
Agent Decision and Verification Data
Extremely strong ↑ New
JEV · Laya · Tev1 · JevBench, and Ollaya emerged in concentration from September 22–25
Computer Use and Cross-System Operations
Strong ↑ New
Qwen/RecreationBench covers 5 operating system categories and 250 application recreation tasks, with 4,404 downloads
Real-World Robot Interaction Data
Strong ↑ New
The NVIDIA Video to Data Challenge covers human demonstration videos, 3D, motion capture, and human-robot interaction
Multimodal 3D and Spatial Data
Strong ↑ New
Meta SHOW3D currently has 20,754 downloads and added synchronized external viewpoints on September 18; UCO3D currently has 112,916 downloads
Speech Intent, Multidialect Support, and Escalation Risk
Strong ↑ New
Google SVQ covers 17 languages; iMerit emphasizes that intent, emotion, and escalation risk matter more than pure transcription
Reward Models and Reward Hacking Data
Strong ↑ New
EleutherAI hack-ignition-benchmark collects RL traces involving exploitable graders; related papers continue to study reward vulnerabilities
Preference Learning and Continuous Feedback
Moderate ↑ New
CCPL studies exoskeleton preferences across environments; MODPO and Off-Policy Evaluation study multi-objective feedback and feedback sampling
Synthetic Boundary Cases and Data Acceptance
Moderate ↑ New
EDGEGEN generates tool-call boundary cases related to enterprise database states; Self-Play Pretraining explores model-generated training data
Scientific RAG Evidence Chains
Moderate ↑ New
The Asta citation counts dataset tracks citations of scientific papers and has 1,403 downloads; Allen AI continues building scientific Agents and VLA evaluation
Physical AI/Embodied Video Data ↓ Dropped Present in previous issue, absent this issue
Code Agent Traces ↓ Dropped Present in previous issue, absent this issue
Reward Hacking and Safety Evaluation Data ↓ Dropped Present in previous issue, absent this issue
Training Data Attribution/Held-Out Query Sets ↓ Dropped Present in previous issue, absent this issue
Legal and Life Sciences Verification Data ↓ Dropped Present in previous issue, absent this issue
Multimodal 3D/Hand-Object Interaction Data ↓ Dropped Present in previous issue, absent this issue
Agent Task Completion Acceptance Data ↓ Dropped Present in previous issue, absent this issue
Medical NLP and Clinical Standard Data ↓ Dropped Present in previous issue, absent this issue
Complex Table Understanding Data ↓ Dropped Present in previous issue, absent this issue
Synthetic Data Quality Evaluation Data ↓ Dropped Present in previous issue, absent this issue

Want to discuss this issue?

Kai
Kai Founder & CEO
苏文
苏文 AI Documentation & Release Engineer
陆明哲
陆明哲 AI Product Manager

Auto-generated by AI Dataset Radar · Updated weekly

AI Dataset Radar →