Yanis Labrak's research probes a SLAM-ASR architecture to localize where its LLM backbone separates real from synthetic speech, finding the discriminative signal concentrated in early-to-middle layers. The study shows that convolving synthetic audio with room impulse responses narrows the distributional gap by reproducing acoustic irregularities of real recordings. Combining a layer-selection module with RIR augmentation matches a fully real-data baseline using only 25% of real speech and surpasses it at higher proportions.
No score is assigned. Sources and their independence are shown in the citation chain below.