Radar Brief Week 27, 2026 · 2026-08-30 — 2026-09-06

BAAI adds three new embodied datasets
Robot collaboration requires human judgment

This week scanned 86 HF orgs · 50 GitHub orgs · 71 blogs · 125 X accounts

0
Valuable Datasets
0
Related Papers
0
Blog Posts
0
Active Repos
One-line Summary

BAAI launched 3 embodied intelligence and collaborative planning datasets on 2026-09-03 [P0], OpenAI simultaneously released GPT-6 Astra and a signal of $1 billion in defensive investment on 2026-09-03 [P0], and AllenAI turned “evaluation data” itself into a product with BenchMIRT on 2026-09-01 [P1]. This week’s strongest data demand signal: embodied intelligence / robot collaboration trajectories.

Key Findings

This week's 5 high commercial value findings

P0 BAAI launched 3 embodied intelligence and collaborative planning datasets on 2026-09-03 [P0]

BAAI/Discoverse-L (2026-09-03, downloads: 129) focuses on EvoVLA visual-language-action research. BAAI/MobileVLA-CoT (2026-09-03, downloads: 56) is a multi-granularity chain-of-thought dataset. BAAI/Orchestra-Bench (2026-09-03, downloads: 186) targets high-level collaborative planning for three robots, with 12,000 samples and multi-view scene inputs.

Business implication → The data focus in embodied intelligence is shifting from “single-task completion” to “multi-robot coordination, cross-view understanding, and verifiable action planning.” This type of data inherently requires human judgment to evaluate whether task decomposition is reasonable, whether collaboration conflicts, and whether actions are safe. For Knowlyr, robot and VLA data will be a high-value category best suited for monetization through “judgment contributions.”
P0 OpenAI simultaneously released GPT-6 Astra and a signal of $1 billion in defensive investment on 2026-09-03 [P0]

OpenAI released GPT-6 Astra (2026-09-03), and officially stated that it reaches next-generation capabilities in computer use, coding, cybersecurity, and science; its safety overview explicitly states that this is the first widely deployed model to reach the Preparedness Framework “Critical” level in cybersecurity capability. On the same day, OpenAI also launched Daybreak for Frontline Defenders and announced a $1 billion investment to support frontline defenders and critical infrastructure.

Business implication → Safety evaluation, red-team adversarial testing, privilege escalation tool selection, and sensitive operation review will continue to grow. The stronger the model, the greater the reliance on human safety experts, because humans still need to judge the boundary between “usable” and “safe to ship.” For Knowlyr, expert judgment in safety, attack-chain review, and task-level risk grading will be clear revenue scenarios.
P1 AllenAI turned “evaluation data” itself into a product with BenchMIRT on 2026-09-01 [P1]

allenai/BenchMIRT (2026-09-01, downloads: 327), allenai/BenchMIRT-item-statistics (2026-09-01, downloads: 352), allenai/BenchMIRT-model-statistics (2026-09-01, downloads: 342), and allenai/BenchMIRT-eval-data (2026-08-25, downloads: 39) all appeared together. AllenAI states that this benchmark is designed to measure LLM latent safety and general reasoning scores, and the data includes prompt-response evaluation records.

Business implication → The industry is shifting from “training answers” to “training how to evaluate answers.” This means demand is rising for scoring rules, preference comparisons, judge consistency, and explanations of failure cases. For Knowlyr, the most valuable asset is not just the answer itself, but the ability to structure expert judgment into evaluation assets.
P1 Code Agent data is shifting from static samples to trajectories and environments: Terminal-Universe and Open-SWE-Traces on 2026-09-03 reinforce each other [P1]

The paper Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments (2026-09-03) directly turns agent trajectories into scalable terminal environments. NVIDIA/Open-SWE-Traces (2026-04-16, downloads: 21,253, likes: 108) continues to update, and its description mentions newly added agent trajectories generated by Qwen3.8-27B. EleutherAI/djinn-problems-v1.0 (2026-08-30, downloads: 95) emphasizes dual-verifier reward-hacking environments, aiming to make the separability of insecure and secure behavior more reliable.

Business implication → Competition in code data has moved from “problem set scale” to “trajectory quality, environment reproducibility, and failure explainability.” What is truly valuable is the execution process that humans can verify, the error paths, and judgment about whether reward hacking exists. This is a classic path for “earning income by contributing judgment.”
P2 Claude Skills, Codex, and the Agent ecosystem continued expanding after 2026-09-03 [P2]

Claude Platform release notes (2026-09-03) v1.30.0 added ant apply, which can create and update agents, environments, skills, memory stores, and deployments from files. On GitHub, anthropics/skills has reached 174,589 stars, anthropics/claude-code 144,185 stars, openai/codex 121,775 stars, openai/openai-agents-python 29,212 stars, and openai/skills 25,462 stars.

Business implication → Agent competition is becoming a “skills asset competition,” not just a prompt competition. In the future, the most needed human contributions will not only be labeling outcomes, but also skill documentation, operating procedures, memory rules, toolchain reviews, and scenario-based evaluation. For Knowlyr, such assets that can orchestrate human knowledge can create a long-term supply moat.

Demand Signals

Infer training data demands from model releases

Data Type Intensity Trend Related Signals
Embodied intelligence / robot collaboration trajectories
extremely strong ↑ New
BAAI launched Discoverse-L · MobileVLA-CoT · Orchestra-Bench in succession; NVIDIA continues to push Open-H Embodiment
Code Agent trajectories / terminal environments
extremely strong ↑ New
Terminal-Universe · Open-SWE-Traces · Claude Code · Codex ecosystem expanding in parallel
Safety evaluation / privilege escalation tool selection
extremely strong ↑ New
GPT-6 Astra’s Critical-level cyber capability; ToolPrivBench; SkillSpector
Evaluation judges / preference judgment data
strong ↑ New
BenchMIRT · RewardBench · Small Language Models as Judges for Rubric-Based RL
Multilingual speech / low-resource languages
strong ↑ New
google/WaxalNLP low-resource African language speech corpus
Video understanding / long-video event boundaries
strong ↑ New
DeepMind agentic video understanding with Gemini
3D / spatial perception / text-to-3D
strong ↑ New
uco3d · Qwen-Drive-1.0-4B · VGGT-Omega
Scientific retrieval / RAG / evidence chains
medium ↑ New
Asta summary citation counts · How AI Is Accelerating Scientific Discovery
Time series forecasting / weather / industrial prediction
medium ↑ New
WeatherNext 3 · timesfm-3.0-pytorch
Robot collaboration / embodied planning data ↓ Dropped Present in previous issue, absent this issue
Tool permissions and safety judgment data ↓ Dropped Present in previous issue, absent this issue
Evaluation and preference statistics data ↓ Dropped Present in previous issue, absent this issue
Agent trajectories and terminal task data ↓ Dropped Present in previous issue, absent this issue
Multilingual speech / TTS data ↓ Dropped Present in previous issue, absent this issue
Multimodal driving / 3D perception data ↓ Dropped Present in previous issue, absent this issue
Preference learning / reward data ↓ Dropped Present in previous issue, absent this issue
Synthetic environments / anti-cheating data ↓ Dropped Present in previous issue, absent this issue
Time series forecasting data ↓ Dropped Present in previous issue, absent this issue

Want to discuss this issue?

Kai
Kai Founder & CEO
苏文
苏文 AI Documentation & Release Engineer
陆明哲
陆明哲 AI Product Manager

Auto-generated by AI Dataset Radar · Updated weekly

AI Dataset Radar →