Adaptive Test-Time Reasoning Distillation workspace. Fine-tuning Nemotron-3-Nano-30B on failure-grounded datasets with PRM-guided GRPO and adaptive budget-forcing.
Identify failure modes and generate synthetic corrections.
Instruction tune on Nemotron formatting tags.
Align steps via Process Reward Model policy updates.
Extend search tokens dynamically for hard problems.