NVIDIA published a tutorial demonstrating how to fine-tune a general-purpose embedding model for domain-specific retrieval using a single GPU in under a day. The pipeline uses synthetic data generation, hard negative mining, and contrastive learning, requiring no manual labeling. Atlassian applied the recipe to its JIRA dataset, improving Recall@60 from 0.751 to 0.951, a 26% increase. NVIDIA reported over 10% improvement in Recall@10 and NDCG@10 on its own documentation dataset.
No score is assigned. Sources and their independence are shown in the citation chain below.