NVIDIA researchers introduced PivotOPD, an on-policy distillation method that trains multi-turn LLM agents to avoid early pivotal mistakes and recover from them, posting the best average against 13 ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results