Applications
REFLEX: Reference-Free Evaluation of Log Summarization via Large Language Model Judgment
arXiv:2511.07458v2 Announce Type: replace Abstract: Evaluating log summarization systems is challenging due to the lack of high-quality reference summaries and the limitations of existing metrics like
arXiv:2511.07458v2 Announce Type: replace Abstract: Evaluating log summarization systems is challenging due to the lack of high-quality reference summaries and the limitations of existing metrics like ROUGE and BLEU, which depend on surface-level lexical overlap. We introduce REFLEX, a reference-free evaluation metric for log summarization based on large language model (LLM) judgment. REFLEX uses LLMs as zero-shot evaluators to assess summary quality along dimensions such as relevance, informativeness, and coherence, without requiring gold-standard references or human annotations. We show that REFLEX produces stable, interpretable, and fine-grained evaluations across multiple log summarization dataset, and more effectively distinguishes model outputs than traditional metrics. REFLEX provides a scalable alternative for evaluating log summaries in real-world settings where reference data is scarce or unavailable.
Related
- FUSE: Ensembling Verifiers with Zero Labeled Data
- PARM: Pipeline-Adapted Reward Model
- RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation
- IUQ: Interrogative Uncertainty Quantification for Long-Form Large Language Model Generation
- Can Large Language Models Detect Methodological Flaws? Evidence from Gesture Recognition for UAV-Based Rescue Operation Based on Deep Learning
Source: arXiv cs.CL | 2026-04-21