BAAI adds three new embodied datasets
Robot collaboration requires human judgment
This week scanned 86 HF orgs · 50 GitHub orgs · 71 blogs · 125 X accounts
BAAI launched 3 embodied intelligence and collaborative planning datasets on 2026-09-03 [P0], OpenAI simultaneously released GPT-6 Astra and a signal of $1 billion in defensive investment on 2026-09-03 [P0], and AllenAI turned “evaluation data” itself into a product with BenchMIRT on 2026-09-01 [P1]. This week’s strongest data demand signal: embodied intelligence / robot collaboration trajectories.
Key Findings
This week's 5 high commercial value findings
BAAI/Discoverse-L (2026-09-03, downloads: 129) focuses on EvoVLA visual-language-action research. BAAI/MobileVLA-CoT (2026-09-03, downloads: 56) is a multi-granularity chain-of-thought dataset. BAAI/Orchestra-Bench (2026-09-03, downloads: 186) targets high-level collaborative planning for three robots, with 12,000 samples and multi-view scene inputs.
OpenAI released GPT-6 Astra (2026-09-03), and officially stated that it reaches next-generation capabilities in computer use, coding, cybersecurity, and science; its safety overview explicitly states that this is the first widely deployed model to reach the Preparedness Framework “Critical” level in cybersecurity capability. On the same day, OpenAI also launched Daybreak for Frontline Defenders and announced a $1 billion investment to support frontline defenders and critical infrastructure.
allenai/BenchMIRT (2026-09-01, downloads: 327), allenai/BenchMIRT-item-statistics (2026-09-01, downloads: 352), allenai/BenchMIRT-model-statistics (2026-09-01, downloads: 342), and allenai/BenchMIRT-eval-data (2026-08-25, downloads: 39) all appeared together. AllenAI states that this benchmark is designed to measure LLM latent safety and general reasoning scores, and the data includes prompt-response evaluation records.
The paper Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments (2026-09-03) directly turns agent trajectories into scalable terminal environments. NVIDIA/Open-SWE-Traces (2026-04-16, downloads: 21,253, likes: 108) continues to update, and its description mentions newly added agent trajectories generated by Qwen3.8-27B. EleutherAI/djinn-problems-v1.0 (2026-08-30, downloads: 95) emphasizes dual-verifier reward-hacking environments, aiming to make the separability of insecure and secure behavior more reliable.
Claude Platform release notes (2026-09-03) v1.30.0 added ant apply, which can create and update agents, environments, skills, memory stores, and deployments from files. On GitHub, anthropics/skills has reached 174,589 stars, anthropics/claude-code 144,185 stars, openai/codex 121,775 stars, openai/openai-agents-python 29,212 stars, and openai/skills 25,462 stars.
Demand Signals
Infer training data demands from model releases
Want to discuss this issue?
Auto-generated by AI Dataset Radar · Updated weekly
AI Dataset Radar →