Liquid AI's LFM2.5-350M scored 22.6% on the IFStruct structured-output benchmark in a local llama.cpp evaluation on an Apple MacBook Pro, close to the reported 21.1%. The accompanying guide shows task-specific fine-tuning of the 350M model in roughly 100 GRPO steps, using about 500 Nemotron structured-output training samples. Prompt augmentation trains code-block formatting and bare-list compliance, aiming to match larger models' performance.
No score is assigned. Sources and their independence are shown in the citation chain below.