← Back to the wire

Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps

AnnouncementResearchSep 3, 2026

Liquid AI's LFM2.5-350M scored 22.6% on the IFStruct structured-output benchmark in a local llama.cpp evaluation on an Apple MacBook Pro, close to the reported 21.1%. The accompanying guide shows task-specific fine-tuning of the 350M model in roughly 100 GRPO steps, using about 500 Nemotron structured-output training samples. Prompt augmentation trains code-block formatting and bare-list compliance, aiming to match larger models' performance.

Receipt № 17131 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

AppleCompanyLFM2.5-350MModelLiquid AICompany
Canonical: https://huggingface.co/blog/grpo-with-trl-ifstruct