AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
3 Jun 2026

Value-Aware Stochastic KV Cache Eviction for Reasoning Models

ResearchDGX agent

arXiv:2606.03928v1 Announce Type: cross Abstract: Reasoning models improve accuracy through extended chains of thought, but their long outputs create a memory and compute bottleneck. KV cache eviction

Visual Instruction Tuning Aligns Modalities through Abstraction

Local AiDGX agent

arXiv:2606.03871v1 Announce Type: cross Abstract: Visual instruction tuning effectively adapts a pre-trained Large Language Model (LLM) to process image information alongside text. Yet, it remains unc

When Does Complexity Conditioning Help a Frozen Sentence Embedding? A Controlled Study of Per-Sentence and Pair-Level Difficulty Adaptation

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.03244v1 Announce Type: new Abstract: A common intuition is that sentence embeddings should adapt to the difficulty of the input. We test this intuition in a controlled, multi-seed setting:

When Models Refuse: Political Steerability and Feature Richness as Measures of Ideological Depth

SafetyDGX agent

arXiv:2508.21448v3 Announce Type: replace Abstract: Large language models (LLMs) sometimes refuse to follow benign instructions, such as declining to argue a political position or adopt a stated perso

Why Are Linear RNNs More Parallelizable?

ResearchDGX agent

arXiv:2603.03612v3 Announce Type: replace-cross Abstract: The community is increasingly exploring linear RNNs (LRNNs) as language models, motivated by their expressive power and parallelizability. Whi

World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

SafetyDGX agent

arXiv:2606.03603v1 Announce Type: cross Abstract: World models and multimodal large language models (MLLMs) provide complementary capabilities for predicting future outcomes from static visual observa

ZX-Calculus:Trace-Indexed Dependent Types and Epistemic Semantics

ResearchDGX agent

arXiv:2606.03063v1 Announce Type: cross Abstract: We propose ZX-Calculus (Knowledge Evolution Calculus), a conservative extension of Martin-Lof Dependent Type Theory (MLTT) integrating trace-indexed t

2 Jun 2026

A Finite-Calibration Regime Map for LLM Judge Panels

Model ReleasesDGX agent

arXiv:2606.01034v1 Announce Type: new Abstract: We study when LLM judge panels should be calibrated with low-dimensional stackers versus joint output tables under finite human-label budgets. Low-dimen

A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL

Model ReleasesDGX agent

arXiv:2606.02398v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves large language models (LLMs) on individual domains such as mathematical reasoning, code generation,

A Registry-Bound LLM Pipeline for Evidence-Grounded Trait Extraction across Tropical Plants, Aquatic Species, and Exotic Pets

ResearchDGX agent

arXiv:2606.00994v1 Announce Type: new Abstract: We describe a registry-bound large-language-model extraction pipeline producing evidence-grounded structured trait records at scale, on cultivated tropi

ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents

Model ReleasesDGX agent

arXiv:2512.00986v3 Announce Type: replace Abstract: A surge in academic publications calls for automated deep research (DR) systems, but accurately evaluating them is still an open problem. First, exi

Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning

AgentsDGX agent

arXiv:2511.14460v2 Announce Type: replace Abstract: Large language models (LLMs) have rapidly evolved from single-turn text generators into the foundation of increasingly capable agents. As these agen

Agentic Clustering: Controllable Text Taxonomies via Multi-Agent Refinement

AgentsDGX agent

arXiv:2606.01255v1 Announce Type: new Abstract: Recent text-clustering methods use large language models to propose a cluster taxonomy from a corpus and then assign each text to it. These pipelines ar

Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why

ResearchDGX agent

arXiv:2606.00093v1 Announce Type: new Abstract: Validating an LLM judge against human annotations usually means reporting several agreement statistics: accuracy, precision, recall, F_1, Cohen's kappa,

AI as a Tool for Simulation-Based Experiments in Literary Studies

ApplicationsDGX agent

arXiv:2606.02293v1 Announce Type: new Abstract: Generative artificial intelligence (AI) systems open new possibilities for experimentation in literary studies via controlled, grounded, large-scale, lo

An Algebraic View of the Expressivity of Recurrent Language Models

ApplicationsDGX agent

arXiv:2606.01765v1 Announce Type: cross Abstract: What formal languages can a recurrent neural language model recognize? Formal results in the literature conflict: some authors report Turing-completen

Are Large Reasoning Models Interruptible?

ApplicationsDGX agent

arXiv:2510.11713v4 Announce Type: replace Abstract: Real-world applications of Large Reasoning Models (LRMs) often require reasoning about changing prompts or environments. In this work, we challenge

ART: Attention Run-time Termination for Efficient Large Language Model Decoding

ResearchDGX agent

arXiv:2606.00024v1 Announce Type: new Abstract: Long-context decoding in Large Language Models (LLMs) is severely constrained by the memory bandwidth required to fetch the extensive Key-Value (KV) cac

Assessment of Generative Named Entity Recognition in the Era of Large Language Models

Model ReleasesDGX agent

arXiv:2601.17898v2 Announce Type: replace Abstract: Named entity recognition (NER) is evolving from a sequence labeling task into a generative paradigm with the rise of large language models (LLMs). W

Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for French

SafetyDGX agent

arXiv:2606.02009v1 Announce Type: new Abstract: In Automated Essay Scoring (AES), benchmarking practices have fostered minimalist evaluation practices, in contrast with the broader-view recommendation

Before and After Temperature: A Distributional View of Creative LLM Generation

Model ReleasesDGX agent

arXiv:2606.01451v1 Announce Type: new Abstract: Reference-free evaluation of large language model (LLM) creativity relies on perplexity, entropy, and top-1 margin. We show that a much stronger signal

Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities

Model ReleasesDGX agent

arXiv:2505.24621v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have transformed natural language understanding and generation, leading to extensive benchmarkin

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

Model ReleasesDGX agent

arXiv:2606.01629v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used for long-form generation, reliably evaluating long-form outputs has become a critical challenge. L

Benchmarking Local LLMs for Natural-Language-to-SQL Querying in Biopharmaceutical Manufacturing: An Empirical Benchmark on Consumer-Grade Hardware

Model ReleasesDGX agent

arXiv:2606.01338v1 Announce Type: new Abstract: Biopharmaceutical manufacturing organizations operate under regulatory frameworks such as FDA guidance, EU Good Manufacturing Practice (GMP), and the EU

Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes

Model ReleasesDGX agent

arXiv:2606.02215v1 Announce Type: new Abstract: Large Language Model (LLM)-augmented Community Notes offer a scalable path for timely, evidence-grounded correction of health misinformation on social p

Beyond Isolated Behaviors: Hierarchical User Modeling for LLM Personalization

Model ReleasesDGX agent

arXiv:2606.02300v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, yet personalizing their outputs to individual users remai

Beyond Scalar Rewards: Dense Feedback for LLM Policy Synthesis in Sequential Social Dilemmas

Model ReleasesDGX agent

arXiv:2603.19453v2 Announce Type: replace Abstract: We study LLM policy synthesis: using a language model to iteratively generate programmatic agent policies for multi-agent environments. Rather than

Beyond Semantic Understanding: Preserving Collaborative Frequency Components in LLM-based Recommendation

Model ReleasesDGX agent

arXiv:2508.10312v2 Announce Type: replace Abstract: Recommender systems in concert with Large Language Models (LLMs) present promising avenues for generating semantically-informed recommendations. How

Beyond Sinusoids: A Morlet Wavelet Framework for Transformer Positional Encoding

ResearchDGX agent

arXiv:2606.01258v1 Announce Type: cross Abstract: Standard positional encodings for transformers - sinusoidal and rotary (RoPE) - treat every position as equally local: they encode where a token is, b

Beyond Topical Similarity: Contrastive Evidence Retrieval with Interpretable Attention Alignment in RAG

SafetyDGX agent

arXiv:2606.01482v1 Announce Type: new Abstract: Ensuring factuality and interpretability in RAG remains an open and urgent problem. We introduce Contrastive Evidence Rationale Attention (CERA), the fi

Beyond Two-Stage Training: Cooperative SFT and RL for LLM Reasoning

ResearchDGX agent

arXiv:2509.06948v3 Announce Type: replace Abstract: Supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) are two widely used post-training paradigms for improving the

BOUTEF: A Multilingual Corpus for FakeNews in North Africa -- Language as a Weapon

ResearchDGX agent

arXiv:2606.00193v1 Announce Type: new Abstract: The rapid spread of fake news on social media has become a major challenge, particularly in multilingual and under-resourced contexts such as North Afri

BranPO: Scalable Contrastive Branch Sampling for Long-Horizon Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2602.03719v2 Announce Type: replace Abstract: Agentic reinforcement learning enables large language models to perform multi-turn planning and tool use, but long-horizon training remains challeng

BraveGuard: From Open-World Threats to Safer Computer-Use Agents

Model ReleasesDGX agent

arXiv:2606.01166v1 Announce Type: cross Abstract: Computer-use agents extend language models from text generation to sustained interaction with files, terminals, browsers, and external tools. This shi

Bridging the Gap: Transfer Learning from English PLMs to Malaysian English

TutorialsDGX agent

arXiv:2407.01374v2 Announce Type: replace Abstract: Malaysian English is a low resource creole language, where it carries the elements of Malay, Chinese, and Tamil languages, in addition to Standard E

Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions

ResearchDGX agent

arXiv:2509.23782v4 Announce Type: replace Abstract: While large language models (LLMs) perform strongly on diverse tasks, their trustworthiness is limited by erratic behavior that is unfaithful to the

CART: Context-Anchored Recurrent Transformer -- A Parameter-Efficient Architecture with Learned Stability

Model ReleasesDGX agent

arXiv:2606.01495v1 Announce Type: cross Abstract: We present CART (Context-Anchored Recurrent Transformer), a parameter-efficient language model that reuses a single shared core block R times across d

CARTE: A Benchmark for Mapping Language Model Knowledge Across France

Model ReleasesDGX agent

arXiv:2606.01995v1 Announce Type: new Abstract: We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language mode

Challenger at MultiPRIDE: Is It Hate Speech or Reclaimed?

ResearchDGX agent

arXiv:2606.01298v1 Announce Type: new Abstract: The spread of hate speech has become increasingly harmful in modern digital environments, particularly on social networking platforms. While recent adva

Characterizing the Effect of Noise in Language Generation in the Limit

ResearchDGX agent

arXiv:2601.21237v2 Announce Type: replace-cross Abstract: Kleinberg and Mullainathan recently proposed a formal framework for studying the phenomenon of language generation, called language generation

Child-directed speech facilitates production, not comprehension, in BabyLMs

Model ReleasesDGX agent

arXiv:2606.01045v1 Announce Type: new Abstract: Recent studies suggest that child-directed speech is not conducive to language learning in BabyLMs. However, current evaluations focus predominantly on

Chunking Methods on Retrieval-Augmented Generation - Effectiveness Evaluation Against Computational Cost and Limitations

ResearchDGX agent

arXiv:2606.00881v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has demonstrated significant capabilities in enhancing the performance of Large Language Models (LLMs). One of the

Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs

Model ReleasesDGX agent

arXiv:2606.00898v1 Announce Type: new Abstract: Large language models systematically hallucinate legal citations -- fabricating statute references, citing repealed provisions, and confusing jurisdicti

ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education

SafetyDGX agent

arXiv:2512.05671v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have achieved remarkable success in dyadic (one-on-one) instruction, they face significant challenges in One-to-M

Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?

AgentsDGX agent

arXiv:2603.03202v3 Announce Type: replace Abstract: As large language models (LLMs) advance their mathematical capabilities toward the IMO and research level, the scarcity of challenging, high-quality

Cognitive-Linguistic Indicators of Depression in Online Communities: Analysed by DistilBERT and Holographic Reduced Representation

ResearchDGX agent

arXiv:2606.00026v1 Announce Type: new Abstract: This paper investigates whether combining cognitively grounded linguistic features with transformer-based embeddings improves automated detection of dep

Confidence-Adaptive SwiGLU for Mixture-of-Experts

ResearchDGX agent

arXiv:2606.00761v1 Announce Type: cross Abstract: SwiGLU has become a standard gated activation in modern Transformer MLPs, yet its gate sharpness -- the smoothness and selectivity of the gating funct

ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?

Model ReleasesDGX agent

arXiv:2606.01849v1 Announce Type: cross Abstract: Differentially private (DP) text synthesis promises to unlock sensitive corpora for model training, but it remains unclear whether DP synthetic data t

Correcting Gradient-Based Circuit Localization via Interaction-Aware Backpropagation

Model ReleasesDGX agent

arXiv:2505.17630v4 Announce Type: replace Abstract: Circuit localization methods aim to identify the subset of model components responsible for specific behaviors in large language models, enabling de

Cost-Aware Diffusion Draft Trees for Speculative Decoding

ResearchDGX agent

arXiv:2606.01813v1 Announce Type: new Abstract: Speculative decoding accelerates inference by having a lightweight drafter propose tokens verified in parallel by the target language model. Block diffu

CRAB-Bench: Evaluating LLM Agents under Complex Task Dependencies and Human-aligned User Simulation

Model ReleasesDGX agent

arXiv:2606.01815v1 Announce Type: new Abstract: Evaluating LLM agents in realistic service scenarios requires complex task dependencies, imperfect user behavior, and an evaluation that accommodates mu

CRAFTQA: A Code-Driven Adaptive Framework for Complex Structured Data Reasoning

TutorialsDGX agent

arXiv:2606.02170v1 Announce Type: new Abstract: Real-world scenarios involve massive heterogeneous structured data (e.g., tables, knowledge graphs), making effective reasoning over such diverse data i

CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning

Model ReleasesDGX agent

arXiv:2606.02502v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) unify heterogeneous vision-language tasks under a shared generative framework via instruction tuning, yet real-

Cross-Environment Neural Reranking for Sample-Efficient Action Selection in Text-Based Agents

Model ReleasesDGX agent

arXiv:2606.02204v1 Announce Type: new Abstract: Large language model agents achieve strong performance on text-based benchmarks but incur prohibitive inference costs, motivating the use of compact neu

Cross-Generational Transfer of Adversarial Attacks Reveals Non-Monotonic Safety Alignment in LLMs

Model ReleasesDGX agent

arXiv:2606.00813v1 Announce Type: cross Abstract: Safety alignment in LLMs does not improve monotonically across model generations. Studying four generations of Google's Gemma family (7B-31B) with qua

Cross-lingual Self-Consistency for Multilingual Reasoning with Language Models

ResearchDGX agent

arXiv:2606.01464v1 Announce Type: new Abstract: Despite expanding their multilingual coverage, the advanced reasoning capabilities of LLMs remain largely confined to a few high-resource languages like

CultureForest: Understanding and Evaluating Cultural Norm Grounded Reasoning in LLMs

Model ReleasesDGX agent

arXiv:2606.01879v1 Announce Type: new Abstract: Existing research largely reduces cultural intelligence in LLMs to a knowledge-level problem, overlooking whether models can effectively utilize their a

CURP: Codebook-based Continuous User Representation for Personalized Generation with LLMs

ResearchDGX agent

arXiv:2602.00742v2 Announce Type: replace Abstract: User modeling characterizes individuals through their preferences and behavioral patterns to enable personalized simulation and generation with Larg

DECK: A Consistency x Confidence Taxonomy of LLM Hallucinations

Model ReleasesDGX agent

arXiv:2606.02289v1 Announce Type: new Abstract: Existing hallucination taxonomies classify LLM errors by what is wrong with the output -- memorised misconceptions, reasoning failures, fluent fabricati

Decoding in Order-Agnostic Language Models: Chain-Rule Deviation and Uniform Spreading

ResearchDGX agent

arXiv:2606.00997v1 Announce Type: new Abstract: Order-agnostic language models (OALMs), including discrete diffusion language models (dLLMs), are trained to predict masked tokens under arbitrary condi

← Previous
1…4748495051…129
Next →