MidTool is a research release for a specific product failure: an agent can recognize a tool yet miss a required argument, lose the workflow, or fail when a user leaves information out. The authors build a 20.3B-token mid-training mix from technical web pages, PDFs, code, real APIs, and Model Context Protocol skills. The release places that training before the usual supervised fine-tuning and reinforcement-learning stages. 1
With the downstream SFT recipe held constant, the paper reports a 4B BFCLv3 overall score of 50.25 after MidTool mid-training, up from 39.73 with SFT alone. The multi-turn average rises from 15.50 to 26.63. The paper also reports a 0.00 web-search subset score in MCP-Universe, so deep-search workflows still need task-specific training evidence. 1
Product teams can start with a shadow evaluation. The team should request access to Arctic-MidTool-RL-4B, pass schemas through Qwen3's
tools= chat template, and replay fixed production-like sessions beside the current route. The team should track schema validity, argument grounding, clarification before a missing input, recovery, p95 latency, and cost. Arctic-MidTool-MT-4B is the checkpoint for teams that will run their own SFT or RL; Arctic-MidTool-RL-4B is the ready-to-run candidate. 23References
- 1
- 2Arctic-MidTool-RL-4B model card
huggingface.co
- 3MidTool-Mix dataset card
huggingface.co


Comments