You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Learned deliberation routing (System III Configurator) from 'Critique of Agent Model' (arXiv:2606.23991): a cost-aware bandit over reasoning depths (direct/cot/plan+verify), live on Vi-GSM8K, with a GRPO path for Qwen3-4B.
Hands-on, runnable demos of SPICED (self-play + a pinch of human data, arxiv 2606.19370) at tiny scale on a Mac M5: demo-regularized self-play, the RLHF KL knob, and LoRA on a real LLM.