← Back to the wire

Validity of LLMs as data annotators: AMALIA on authority

AchievementResearchJul 10, 2026

AMALIA, a 9B-parameter European Portuguese model, agrees with human coders within six F1 points of models eight to thirteen times larger. However, researcher Manuel Pita found that when holistic prompts were decomposed into atomic clauses, AMALIA recovered only about half its performance, suggesting reliance on surface correlates. A calibrated English instrument did not transfer to AMALIA. The study argues sovereign-LLM benchmarks should test not only agreement with human coders but the evidential route warranting that agreement.

Receipt № 4721 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

01medPRIMARY
AMALIAModelManuel PitaPerson
Canonical: https://arxiv.org/abs/2607.08731