U2 Flash:为真实生产任务而生
U2 Flash 聚焦Agent 与 Coding核心场景,具备强大的复杂推理、代码生成与任务执行能力,支持长上下文,可胜任复杂工程开发、长程任务及多步骤Agent 工作流。在保持强劲模型能力的同时,实现更快的推理速度与更优的成本效率,为开发者与企业提供高性能、高性价比的智能底座。
Highlights
High-density intelligence base
A 300B-scale model with efficient parameter use for long-horizon tasks and engineering-grade coding, balancing capability and deployment cost.
Long-horizon task capability
When facing complex tasks and blockers, it tries alternative paths and keeps searching for workable solutions so tasks run through more reliably.
Autonomous closed-loop evolution
Self-evolving optimization supports autonomous correction and iterative tuning, reducing heavy human labeling while continuously upgrading the model.
Autonomous closed-loop evolution
Dual-track data generation
For a batch of tasks, multiple strong teacher models produce correct trajectories while a weak model produces incorrect ones, forming paired right/wrong training samples.
Supervised analysis & key-action marking
A supervisor model with harness compares paired trajectories, produces analysis reports and flowcharts, and marks decisive actions.
Prompted CoT generation
From key-action marks, generate prompts and chain-of-thought to focus the weak model on its weak spots.
Weak-model re-sampling
The weak model re-samples and produces correct trajectories guided by prompted CoT.
Cleaning & rewrite
Clean and rewrite prompts and reasoning traces to improve data quality.
Fine-tune & reinforce (SFT + async Agent RL + OPD)
SFT on cleaned data, then async Agent RL in a real sandbox, then multi-teacher OPD to distill specialist skills back into one model; the enhanced model becomes the next-round weak model, closing the loop.
Self-strengthening in real environments
Async sampling
Multiple rollout workers run tasks and collect trajectories (code execution, tool calls, environment interaction) without blocking policy updates, boosting throughput.
GRPO policy optimization
Following DeepSeekMath’s GRPO (Group Relative Policy Optimization), relative advantages are estimated within a trajectory group—no separate Critic—cutting memory and cost while diverse in-group paths encourage multi-solution exploration.
Sparse reward + long-horizon credit
Agent rewards are sparse; key-action marks and trajectory-level advantages back-propagate success/failure to decisive actions, addressing long-horizon credit assignment.
Supervisor only at key steps
Aligned with the closed loop—the model explores freely while the supervisor intervenes only at critical forks, balancing exploration and quality.
Multi-teacher on-policy distillation
Domain experts
Train an expert per domain (math / code / Agent / instruction following) with SFT plus domain GRPO to form an expert pool.
Multi-teacher merge
Distill N experts into one student that minimizes reverse KL to each teacher on its own on-policy trajectories.
Engineering
Load teacher weights on demand, cache hidden states only, and order samples by teacher—multi-teacher distillation under limited compute.