AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
CompaniesToolsTechniques

Each lane shows up to 8 recent matching entries, ordered from earlier to later. Tracks load separately to keep the 75,000+ entry wiki fast.

Companies

CompanyAnthropic8 recent entries
7 Aug 2026Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge Evaluation

arXiv:2608.05353v1 Announce Type: new Abstract: LLM judges are often asked to extract criteria and evidence before choosing between candidate answers. This workflow assumes that the intermediate recor

→10 Aug 2026Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence

arXiv:2608.06778v1 Announce Type: cross Abstract: Mapping cyber threat intelligence (CTI) text to MITRE ATT&CK techniques is essential for structured threat analysis, yet manual annotation is costly a

3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
→10 Aug 2026GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization

arXiv:2608.06526v1 Announce Type: new Abstract: Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into

→11 Aug 2026DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?

arXiv:2608.07614v1 Announce Type: cross Abstract: Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code

→11 Aug 2026BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

arXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi

→12 Aug 2026VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world

→12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;

→12 Aug 2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses

arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these

CompanyOpenAI8 recent entries
31 Jul 2026Generative AI and linguistic diversity in academic writing and publishing: Perspectives from World Englishes

arXiv:2607.28505v1 Announce Type: new Abstract: The rise of generative artificial intelligence (GenAI) in academic writing and publishing (AWP) raises questions about linguistic inclusivity and the le

→5 Aug 2026Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

→5 Aug 2026Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m

→5 Aug 2026ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference

arXiv:2608.02947v1 Announce Type: cross Abstract: The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelen

→6 Aug 2026Document Optimization for Black-Box Retrieval via Reinforcement Learning

arXiv:2604.05087v3 Announce Type: replace Abstract: Document expansion is a classical technique for improving retrieval quality, and is attractive since it shifts computation offline, avoiding additio

→10 Aug 2026Grammar Engineering Meets LLMs: Development of Cantonese and Irish ParGram Treebanks

arXiv:2608.07283v1 Announce Type: new Abstract: Grammar engineering requires expertise in linguistic formalism and computational implementation, especially in parallel grammar projects that balance cr

→11 Aug 2026Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasks

arXiv:2512.22966v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision

→11 Aug 2026DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?

arXiv:2608.07614v1 Announce Type: cross Abstract: Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code

CompanyGoogle8 recent entries
6 Aug 2026Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness

arXiv:2607.19322v2 Announce Type: replace Abstract: Rubric-based evaluation of open-ended generation faces a fundamental tension between expressiveness and reliability. Authoring a faithful rubric req

→6 Aug 2026Transfer Learning for Named Entity Recognition of Classical Latin through LLM Prompting

arXiv:2608.04015v1 Announce Type: new Abstract: With the increase in digitized resources of Classical Latin texts and modern breakthroughs of Large Language Models (LLMs), I contribute to ancient lang

→10 Aug 2026GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization

arXiv:2608.06526v1 Announce Type: new Abstract: Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into

→10 Aug 2026Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation

arXiv:2608.06718v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paraling

→11 Aug 2026The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction

arXiv:2608.07517v1 Announce Type: cross Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aw

→11 Aug 2026Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasks

arXiv:2512.22966v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision

→11 Aug 2026LexKairos: Benchmarking Legal Temporal Capabilities in LLMs

arXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that

→11 Aug 2026BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

arXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi

CompanyMeta8 recent entries
10 Aug 2026Latent Fact-Checking: Detecting Misinformation through Activation Engineering

arXiv:2608.06417v1 Announce Type: cross Abstract: The proliferation of misinformation online has driven demand for scalable detection systems. While most existing approaches rely on surface-level ling

→10 Aug 2026GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization

arXiv:2608.06526v1 Announce Type: new Abstract: Large language models can infer sensitive personal attributes, such as age, location, and occupation, from ordinary text, turning everyday writing into

→11 Aug 2026SciTaRC: A Plan-Annotated Scientific Tabular QA Benchmark for Language Reasoning and Complex Computation

arXiv:2603.08910v2 Announce Type: replace Abstract: We introduce SciTaRC, an expert-authored benchmark for question answering over scientific tables that targets composite, multi-step reasoning. To en

→11 Aug 2026SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems

arXiv:2608.08237v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cos

→11 Aug 2026Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages

arXiv:2608.08800v1 Announce Type: new Abstract: Pretraining LLMs on artificial languages ('pre-pretraining') is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33%

→11 Aug 2026Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation

arXiv:2608.09140v1 Announce Type: cross Abstract: Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information

→12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;

→12 Aug 2026Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

arXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that th

CompanyMistral8 recent entries
5 Aug 2026Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

arXiv:2608.03842v1 Announce Type: new Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), 'which layer is responsible' has three natural operationalization

→7 Aug 2026PoolBench: A Benchmark for Pooling Strategies in Concept Representation Evaluation for Decoder-Only LLMs

arXiv:2608.05162v1 Announce Type: new Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden

→7 Aug 2026Human-Like Anaphor Resolution in Large Language Models

arXiv:2608.05630v1 Announce Type: new Abstract: Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science

→11 Aug 2026SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems

arXiv:2608.08237v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cos

→11 Aug 2026Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

arXiv:2608.09046v1 Announce Type: new Abstract: Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure doe

→11 Aug 2026High-Layer Attention Pruning with Rescaling

arXiv:2507.01900v3 Announce Type: replace Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional

→12 Aug 2026REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs

arXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud

→12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;

CompanyxAI8 recent entries
24 Jul 2026Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks

arXiv:2607.20864v1 Announce Type: cross Abstract: Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single ans

→27 Jul 2026Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

arXiv:2607.22513v1 Announce Type: cross Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor

→28 Jul 2026Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It

arXiv:2607.23893v1 Announce Type: cross Abstract: Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the questio

→28 Jul 2026Tailored untruths: How personalisation challenges LLM safeguards

arXiv:2510.12993v3 Announce Type: replace Abstract: Large Language Models (LLMs) can generate highly persuasive disinformation, yet little is known about how effectively they personalise it across lan

→28 Jul 2026Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration

arXiv:2607.23538v1 Announce Type: new Abstract: Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languag

→31 Jul 2026Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabil

→4 Aug 2026When Retrieval Helps and Distracts: Evaluating Evidence-Generating LLMs for Biomedical Claim Verification

arXiv:2608.01409v1 Announce Type: new Abstract: Biomedical fact-checking systems must do more than predict whether a claim is supported, contradicted, or unaddressed: they should also produce evidence

→4 Aug 2026From We to Me: Theory Informed Narrative Shift with Abductive Reasoning

arXiv:2603.03320v2 Announce Type: replace Abstract: Effective communication often relies on aligning a message with an audience's narrative and worldview. Narrative shift involves transforming text to

CompanyDeepSeek8 recent entries
4 Aug 2026From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States

arXiv:2607.09842v2 Announce Type: replace-cross Abstract: We investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the hidden-state tra

→4 Aug 2026Cost-Effective Automated Judging of Natural-Language Mathematical Proofs

arXiv:2608.00004v1 Announce Type: new Abstract: Grading natural-language mathematical proofs is a recurring cost in evaluating math-reasoning systems, and frontier LLM judges are expensive. We ask whe

→5 Aug 2026SeqLLM: Augmenting LLMs with Behavioral-Sequence Modeling for High-Stakes Decisions at WeChat Pay

arXiv:2608.03063v1 Announce Type: new Abstract: Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false

→5 Aug 2026Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

→6 Aug 2026Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?

arXiv:2608.05097v1 Announce Type: new Abstract: Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same

→7 Aug 2026ConWriter: Transition-Constrained Stateful Long-Form Story Generation with Lightweight Neuro-Symbolic Consistency Control

arXiv:2608.05169v1 Announce Type: new Abstract: Long-form story generation requires models to preserve narrative consistency across extended contexts, yet existing prompting-based methods often accumu

→11 Aug 2026Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving

arXiv:2608.08910v1 Announce Type: new Abstract: PTQTP decomposes LLM weight matrices into two ternary (trit) planes with two free per-group scales. Tying the scales to a fixed ratio of three collapses

→11 Aug 2026Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders

arXiv:2608.08168v1 Announce Type: new Abstract: While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this e

CompanyNVIDIA8 recent entries
4 Aug 2026Writing-System-Level Tokenizer Adaptation for Byte-Level BPE

arXiv:2608.00582v1 Announce Type: new Abstract: Pretrained byte-level BPE tokenizers can segment underrepresented languages inefficiently. Replacing a tokenizer changes the meaning of nearly every tok

→4 Aug 2026RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation

arXiv:2608.01810v1 Announce Type: new Abstract: Rubric-based LLM-as-judge pipelines often assume that evaluation criteria provide independent signals. In practice, however, criteria can be behaviorall

→4 Aug 2026DiffusionGemma Technical Report

arXiv:2608.00146v1 Announce Type: new Abstract: We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rathe

→5 Aug 2026Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension

arXiv:2608.03494v1 Announce Type: new Abstract: Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token

→6 Aug 2026The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale

arXiv:2608.04355v1 Announce Type: new Abstract: Accuracy changes after language-model self-revision are usually interpreted as changes in reasoning. We show this can fail at the answer-extraction boun

→10 Aug 2026Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control

arXiv:2608.06758v1 Announce Type: new Abstract: We present Stockmark-Nemotron-3-Nano-Omni-JapanDocReader, a Japanese document understanding model built from Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16

→11 Aug 2026Memorization Dynamics in Knowledge Distillation for Language Models

arXiv:2601.15394v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is increasingly adopted to transfer capabilities from large language models to smaller ones, offering significant improv

→11 Aug 2026AraSSM: A bidirectional state-space encoder for Arabic masked language modeling

arXiv:2608.08256v1 Announce Type: new Abstract: Pretrained Transformer encoders such as AraBERT, MARBERT, and CAMeLBERT have become the standard backbone for Arabic natural language understanding, but