Model Releases
Evaluating RAG for French immigration law: a benchmark and baseline study
arXiv:2607.24449v1 Announce Type: cross Abstract: International recruitment in France requires navigating a layered legal framework absent from existing legal AI benchmarks. We present a publicly avai
arXiv:2607.24449v1 Announce Type: cross Abstract: International recruitment in France requires navigating a layered legal framework absent from existing legal AI benchmarks. We present a publicly available benchmark and first comparative evaluation for this domain, covering permit-type recommendation, required-document retrieval, and legal citation coverage. Comparing a parametric LLM baseline against dense retrieval augmentation at two model scales (Qwen3.5-9B and -27B) on 52 annotated synthetic profiles, we find that retrieval improves administrative guidance at both scales, most notably permit-type accuracy. Our results confirm that retrieval grounding is important for more reliable administrative guidance in this domain, and motivate further investigation of hybrid retrieval strategies.
Related
- UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
- Magis-Bench: Evaluating LLMs on Magistrate-Level Legal Tasks
- Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
Source: arXiv cs.AI | 2026-07-28