CompanyAnthropic8 recent entries7 Aug 2026Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge EvaluationarXiv:2608.05353v1 Announce Type: new Abstract: LLM judges are often asked to extract criteria and evidence before choosing between candidate answers. This workflow assumes that the intermediate recor→10 Aug 2026Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat IntelligencearXiv:2608.06778v1 Announce Type: cross Abstract: Mapping cyber threat intelligence (CTI) text to MITRE ATT&CK techniques is essential for structured threat analysis, yet manual annotation is costly a
CompanyOpenAI8 recent entries31 Jul 2026Generative AI and linguistic diversity in academic writing and publishing: Perspectives from World EnglishesarXiv:2607.28505v1 Announce Type: new Abstract: The rise of generative artificial intelligence (GenAI) in academic writing and publishing (AWP) raises questions about linguistic inclusivity and the le→5 Aug 2026Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI SafetyarXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive→5 Aug 2026Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language ModelsarXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m→5 Aug 2026ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM InferencearXiv:2608.02947v1 Announce Type: cross Abstract: The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelen→6 Aug 2026Document Optimization for Black-Box Retrieval via Reinforcement LearningarXiv:2604.05087v3 Announce Type: replace Abstract: Document expansion is a classical technique for improving retrieval quality, and is attractive since it shifts computation offline, avoiding additio→10 Aug 2026Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram TreebanksarXiv:2608.07283v1 Announce Type: new Abstract: Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cr→11 Aug 2026Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasksarXiv:2512.22966v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision→11 Aug 2026DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?arXiv:2608.07614v1 Announce Type: cross Abstract: Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code
CompanyGoogle8 recent entries6 Aug 2026Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual CompletenessarXiv:2607.19322v2 Announce Type: replace Abstract: Rubric-based evaluation of open-ended generation faces a fundamental tension between expressiveness and reliability. Authoring a faithful rubric req→6 Aug 2026Transfer Learning for Named Entity Recognition of Classical Latin through LLM PromptingarXiv:2608.04015v1 Announce Type: new Abstract: With the increase in digitized resources of Classical Latin texts and modern breakthroughs of Large Language Models (LLMs), I contribute to ancient lang→10 Aug 2026GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy OptimizationarXiv:2608.06526v1 Announce Type: new Abstract: Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into →10 Aug 2026Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response EvaluationarXiv:2608.06718v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paraling→11 Aug 2026The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test PredictionarXiv:2608.07517v1 Announce Type: cross Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aw→11 Aug 2026Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasksarXiv:2512.22966v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision→11 Aug 2026LexKairos: Benchmarking Legal Temporal Capabilities in LLMsarXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that→11 Aug 2026BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and MitigationarXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi
CompanyMeta8 recent entries10 Aug 2026Latent Fact-Checking: Detecting Misinformation through Activation EngineeringarXiv:2608.06417v1 Announce Type: cross Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level ling→10 Aug 2026GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy OptimizationarXiv:2608.06526v1 Announce Type: new Abstract: Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into →11 Aug 2026SciTaRC: A Plan-Annotated Scientific Tabular QA Benchmark for Language Reasoning and Complex ComputationarXiv:2603.08910v2 Announce Type: replace Abstract: We introduce SciTaRC, an expert-authored benchmark for question answering over scientific tables that targets composite, multi-step reasoning. To en→11 Aug 2026SAGE: SLO-Aware Adaptive Retrieval for Production RAG SystemsarXiv:2608.08237v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cos→11 Aug 2026Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple LanguagesarXiv:2608.08800v1 Announce Type: new Abstract: Pretraining LLMs on artificial languages ('pre-pretraining') is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33%→11 Aug 2026Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query DelegationarXiv:2608.09140v1 Announce Type: cross Abstract: Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information →12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog FaithfulnessarXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;→12 Aug 2026Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context ExtensionarXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that th
CompanyMistral8 recent entries5 Aug 2026Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its ScalingarXiv:2608.03842v1 Announce Type: new Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), 'which layer is responsible' has three natural operationalization→7 Aug 2026PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMsarXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden →7 Aug 2026Human-Like Anaphor Resolution in Large Language ModelsarXiv:2608.05630v1 Announce Type: new Abstract: Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science →11 Aug 2026SAGE: SLO-Aware Adaptive Retrieval for Production RAG SystemsarXiv:2608.08237v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cos→11 Aug 2026Measuring the Tokenization Premium: A Cost Audit for Underserved Language CommunitiesarXiv:2608.09046v1 Announce Type: new Abstract: Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure doe→11 Aug 2026High-Layer Attention Pruning with RescalingarXiv:2507.01900v3 Announce Type: replace Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional→12 Aug 2026REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMsarXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud→12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog FaithfulnessarXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;
CompanyxAI8 recent entries24 Jul 2026Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM BenchmarksarXiv:2607.20864v1 Announce Type: cross Abstract: Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single ans→27 Jul 2026Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-SciencearXiv:2607.22513v1 Announce Type: cross Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor →28 Jul 2026Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of ItarXiv:2607.23893v1 Announce Type: cross Abstract: Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the questio→28 Jul 2026Tailored untruths: How personalisation challenges LLM safeguardsarXiv:2510.12993v3 Announce Type: replace Abstract: Large Language Models (LLMs) can generate highly persuasive disinformation, yet little is known about how effectively they personalise it across lan→28 Jul 2026Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM CollaborationarXiv:2607.23538v1 Announce Type: new Abstract: Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languag→31 Jul 2026Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction GamearXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabil→4 Aug 2026When Retrieval Helps and Distracts: Evaluating Evidence-Generating LLMs for Biomedical Claim VerificationarXiv:2608.01409v1 Announce Type: new Abstract: Biomedical fact-checking systems must do more than predict whether a claim is supported, contradicted, or unaddressed: they should also produce evidence→4 Aug 2026From We to Me: Theory Informed Narrative Shift with Abductive ReasoningarXiv:2603.03320v2 Announce Type: replace Abstract: Effective communication often relies on aligning a message with an audience's narrative and worldview. Narrative shift involves transforming text to
CompanyDeepSeek8 recent entries4 Aug 2026From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden StatesarXiv:2607.09842v2 Announce Type: replace-cross Abstract: We investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the hidden-state tra→4 Aug 2026Cost-Effective Automated Judging of Natural-Language Mathematical ProofsarXiv:2608.00004v1 Announce Type: new Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whe→5 Aug 2026SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat PayarXiv:2608.03063v1 Announce Type: new Abstract: Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false →5 Aug 2026Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI SafetyarXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive→6 Aug 2026Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?arXiv:2608.05097v1 Announce Type: new Abstract: Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same →7 Aug 2026ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency ControlarXiv:2608.05169v1 Announce Type: new Abstract: Long-form story generation requires models to preserve narrative consistency across extended contexts, yet existing prompting-based methods often accumu→11 Aug 2026Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts ServingarXiv:2608.08910v1 Announce Type: new Abstract: PTQTP decomposes LLM weight matrices into two ternary (trit) planes with two free per-group scales. Tying the scales to a fixed ratio of three collapses→11 Aug 2026Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse AutoencodersarXiv:2608.08168v1 Announce Type: new Abstract: While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this e
CompanyNVIDIA8 recent entries4 Aug 2026Writing-System-Level Tokenizer Adaptation for Byte-Level BPEarXiv:2608.00582v1 Announce Type: new Abstract: Pretrained byte-level BPE tokenizers can segment underrepresented languages inefficiently. Replacing a tokenizer changes the meaning of nearly every tok→4 Aug 2026RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge EvaluationarXiv:2608.01810v1 Announce Type: new Abstract: Rubric-based LLM-as-judge pipelines often assume that evaluation criteria provide independent signals. In practice, however, criteria can be behaviorall→4 Aug 2026DiffusionGemma Technical ReportarXiv:2608.00146v1 Announce Type: new Abstract: We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rathe→5 Aug 2026Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary ExtensionarXiv:2608.03494v1 Announce Type: new Abstract: Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token →6 Aug 2026The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid ScalearXiv:2608.04355v1 Announce Type: new Abstract: Accuracy changes after language-model self-revision are usually interpreted as changes in reasoning. We show this can fail at the answer-extraction boun→10 Aug 2026Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting ControlarXiv:2608.06758v1 Announce Type: new Abstract: We present Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built from Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16→11 Aug 2026Memorization Dynamics in Knowledge Distillation for Language ModelsarXiv:2601.15394v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is increasingly adopted to transfer capabilities from large language models to smaller ones, offering significant improv→11 Aug 2026AraSSM: A bidirectional state-space encoder for Arabic masked language modelingarXiv:2608.08256v1 Announce Type: new Abstract: Pretrained Transformer encoders such as AraBERT, MARBERT, and CAMeLBERT have become the standard backbone for Arabic natural language understanding, but