AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
21 Apr 2026

GenericAgent: A Token-Efficient Self-Evolving LLM Agent via Contextual Information Density Maximization (V1.0)

AgentsDGX agent

arXiv:2604.17091v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are fundamentally limited by context. As interactions become longer, tool descriptions, retrieved memorie

Geometric Stability: The Missing Axis of Representations

SafetyDGX agent

arXiv:2601.09173v4 Announce Type: replace-cross Abstract: Representational similarity analysis and related methods have become standard tools for comparing the internal geometries of neural networks a

GeometryZero: Advancing Geometry Solving via Group Contrastive Policy Optimization

SafetyDGX agent

arXiv:2506.07160v3 Announce Type: replace Abstract: Recent progress in large language models (LLMs) has boosted mathematical reasoning, yet geometry remains challenging where auxiliary construction is


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

GeoRC: A Benchmark for Geolocation Reasoning Chains

Model ReleasesDGX agent

arXiv:2601.21278v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) are good at recognizing the global location of a photograph -- their geolocation prediction accuracy rivals the

GoCoMA: Hyperbolic Multimodal Representation Fusion for Large Language Model-Generated Code Attribution

ResearchDGX agent

arXiv:2604.16377v1 Announce Type: new Abstract: Large Language Models (LLMs) trained on massive code corpora are now increasingly capable of generating code that is hard to distinguish from human-writ

GraSP: Graph-Structured Skill Compositions for LLM Agents

Local AiDGX agent

arXiv:2604.17870v1 Announce Type: new Abstract: Skill ecosystems for LLM agents have matured rapidly, yet recent benchmarks show that providing agents with more skills does not monotonically improve p

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling

Model ReleasesDGX agent

arXiv:2604.18556v1 Announce Type: new Abstract: Weight quantization has become a standard tool for efficient LLM deployment, especially for local inference, where models are now routinely served at 2-

HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders

Model ReleasesDGX agent

arXiv:2604.16430v1 Announce Type: new Abstract: Large Language Models (LLMs) are powerful and widely adopted, but their practical impact is limited by the well-known hallucination phenomenon. While re

Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages

ResearchDGX agent

arXiv:2604.18204v1 Announce Type: new Abstract: We present a phoneme-level analysis of automatic speech recognition (ASR) for two low-resourced and phonologically complex East Caucasian languages, Arc

HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents

AgentsDGX agent

arXiv:2604.16839v1 Announce Type: new Abstract: Long-term memory is a critical challenge for Large Language Model agents, as fixed context windows cannot preserve coherence across extended interaction

HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference

ResearchDGX agent

arXiv:2601.13684v2 Announce Type: replace Abstract: The linear memory growth of the KV cache poses a significant bottleneck for LLM inference in long-context tasks. Existing static compression methods

Heterogeneity in Formal Linguistic Competence of Language Models: Is Data the Real Bottleneck?

TutorialsDGX agent

arXiv:2604.17930v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit a puzzling disparity in their formal linguistic competence: while they learn some linguistic phenomena with near-pe

Hierarchical Retrieval with Out-Of-Vocabulary Queries: A Case Study on SNOMED CT

ApplicationsDGX agent

arXiv:2511.16698v2 Announce Type: replace Abstract: SNOMED CT is a biomedical ontology with a hierarchical representation, modelling terminological concepts at a large scale. Knowledge retrieval in SN

HiGMem: A Hierarchical and LLM-Guided Memory System for Long-Term Conversational Agents

Model ReleasesDGX agent

arXiv:2604.18349v1 Announce Type: new Abstract: Long-term conversational large language model (LLM) agents require memory systems that can recover relevant evidence from historical interactions withou

HiP-LoRA: Budgeted Spectral Plasticity for Robust Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2604.17751v1 Announce Type: cross Abstract: Adapting foundation models under resource budgets relies heavily on Parameter-Efficient Fine-Tuning (PEFT), with LoRA being a standard modular solutio

HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution

Model ReleasesDGX agent

arXiv:2604.17745v1 Announce Type: new Abstract: Recent advances in large language models have highlighted their potential to automate computational research, particularly reproducing experimental resu

HopRank: Self-Supervised LLM Preference-Tuning on Graphs for Few-Shot Node Classification

ApplicationsDGX agent

arXiv:2604.17271v1 Announce Type: new Abstract: Node classification on text-attributed graphs (TAGs) is a fundamental task with broad applications in citation analysis, social networks, and recommenda

HopWeaver: Cross-Document Synthesis of High-Quality and Authentic Multi-Hop Questions

ResearchDGX agent

arXiv:2505.15087v3 Announce Type: replace Abstract: Multi-Hop Question Answering (MHQA) is crucial for evaluating the model's capability to integrate information from diverse sources. However, creatin

HORIZON: A Benchmark for In-the-wild User Behaviour Modeling

Model ReleasesDGX agent

arXiv:2604.17259v1 Announce Type: cross Abstract: User behavior in the real world is diverse, cross-domain, and spans long time horizons. Existing user modeling benchmarks however remain narrow, focus

HorizonBench: Long-Horizon Personalization with Evolving Preferences

Model ReleasesDGX agent

arXiv:2604.17283v1 Announce Type: new Abstract: User preferences evolve across months of interaction, and tracking them requires inferring when a stated preference has been changed by a subsequent lif

How Creative Are Large Language Models in Generating Molecules?

Local AiDGX agent

arXiv:2604.18031v1 Announce Type: new Abstract: Molecule generation requires satisfying multiple chemical and biological constraints while searching a large and structured chemical space. This makes i

How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effects

SafetyDGX agent

arXiv:2510.06700v3 Announce Type: replace Abstract: Both humans and large language models (LLMs) exhibit content effects: biases in which the plausibility of the semantic content of a reasoning proble

How Non-Linguistic Is the Indus Sign System? A Synthetic-Baseline Scorecard

ApplicationsDGX agent

arXiv:2604.17828v1 Announce Type: new Abstract: Whether the Indus Valley sign system (c. 2600-1900 BCE) encodes spoken language has been debated for decades. This paper introduces a multi-metric discr

How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study

Model ReleasesDGX agent

arXiv:2505.15404v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have achieved remarkable success on reasoning-intensive tasks such as mathematics and programming. However, their enha

How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve Them

Local AiDGX agent

arXiv:2604.17105v1 Announce Type: new Abstract: Tokenization is the first step in every language model (LM), yet it never takes the sounds of words into account. We investigate how tokenization influe

How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models

ApplicationsDGX agent

arXiv:2510.02370v3 Announce Type: replace Abstract: Large language models leverage both parametric knowledge acquired during pretraining and in-context knowledge provided at inference time. Crucially,

HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models

ResearchDGX agent

arXiv:2511.01066v3 Announce Type: replace Abstract: We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 tr

Human-Centered Supervision for Sentiment Analysis in Telugu: A Systematic Inquiry Beyond Accuracy

SafetyDGX agent

arXiv:2508.01486v3 Announce Type: replace Abstract: Sentiment analysis for low-resource languages remains challenging in an era where interpretability, human alignment, and fairness are increasingly n

IceBreaker for Conversational Agents: Breaking the First-Message Barrier with Personalized Starters

SafetyDGX agent

arXiv:2604.18375v1 Announce Type: new Abstract: Conversational agents, such as ChatGPT and Doubao, have become essential daily assistants for billions of users. To further enhance engagement, these sy

ICLAD: In-Context Learning with Comparison-Guidance for Audio Deepfake Detection

ResearchDGX agent

arXiv:2604.16749v1 Announce Type: cross Abstract: Audio deepfakes pose a significant security threat, yet current state-of-the-art (SOTA) detection systems do not generalize well to realistic in-the-w

ImpRIF: Stronger Implicit Reasoning Leads to Better Complex Instruction Following

ResearchDGX agent

arXiv:2602.21228v2 Announce Type: replace Abstract: As applications of large language models (LLMs) become increasingly complex, the demand for robust complex instruction following capabilities is gro

Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification

Model ReleasesDGX agent

arXiv:2604.17010v1 Announce Type: new Abstract: We introduce a self-play framework for semantic equivalence in Haskell, utilizing formal verification to guide adversarial training between a generator

Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context

ResearchDGX agent

arXiv:2506.10779v2 Announce Type: replace Abstract: Classroom speech and lectures often contain named entities (NEs) such as names of people and special terminology. While automatic speech recognition

Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation

ResearchDGX agent

arXiv:2604.16881v1 Announce Type: new Abstract: Cross-cultural entity translation remains challenging for large language models (LLMs) as literal or phonetic renderings are usually yielded instead of

Inertia in Moral and Value Judgments of Large Language Models

SafetyDGX agent

arXiv:2408.09049v3 Announce Type: replace Abstract: Large Language Models (LLMs) behave non-deterministically, and prompting has become a common method for steering their outputs. A popular strategy i

Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluation

Model ReleasesDGX agent

arXiv:2510.09275v2 Announce Type: replace Abstract: Medical diagnostics is a high-stakes and complex domain that is critical to patient care. However, current evaluations of large language models (LLM

Information Representation Fairness in Long-Document Embeddings: The Peculiar Interaction of Positional and Language Bias

SafetyDGX agent

arXiv:2601.16934v2 Announce Type: replace Abstract: To be discoverable in an embedding-based search process, each part of a document should be reflected in its embedding representation. To quantify an

Injecting Structured Biomedical Knowledge into Language Models: Continual Pretraining vs. GraphRAG

Model ReleasesDGX agent

arXiv:2604.16422v1 Announce Type: new Abstract: The injection of domain-specific knowledge is crucial for adapting language models (LMs) to specialized fields such as biomedicine. While most current a

iPhoneme: Brain-to-Text Communication for ALS Using ConformerXL Decoding

ResearchDGX agent

arXiv:2604.16441v1 Announce Type: cross Abstract: Brain-computer interfaces (BCIs) for speech restoration hold transformative potential for the approximately 173,000--232,500 individuals worldwide wit

Is Agentic RAG worth it? An experimental comparison of RAG approaches

AgentsDGX agent

arXiv:2601.07711v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems are usually defined by the combination of a generator and a retrieval component that extracts textual c

IYKYK (But AI Doesn't): Automated Content Moderation Does Not Capture Communities' Heterogeneous Attitudes Towards Reclaimed Language

SafetyDGX agent

arXiv:2604.16654v1 Announce Type: new Abstract: Reclaimed slur usage is a common and meaningful practice online for many marginalized communities. It serves as a source of solidarity, identity, and sh

Jailbreaking Large Language Models with Morality Attacks

SafetyDGX agent

arXiv:2604.17053v1 Announce Type: new Abstract: Pluralism alignment with AI has the sophisticated and necessary goal of creating AI that can coexist with and serve morally multifaceted humanity. Resea

JudgeMeNot: Personalizing Large Language Models to Emulate Judicial Reasoning in Hebrew

Model ReleasesDGX agent

arXiv:2604.18041v1 Announce Type: new Abstract: Despite significant advances in large language models, personalizing them for individual decision-makers remains an open problem. Here, we introduce a s

Jupiter-N Technical Report

Model ReleasesDGX agent

arXiv:2604.17429v1 Announce Type: new Abstract: We present Jupiter-N, a hybrid reasoning model post-trained from Nemotron 3 Super, a fully open-source 120 billion parameter LLM. We target three object

Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning

Model ReleasesDGX agent

arXiv:2604.18419v1 Announce Type: cross Abstract: Large language models (LLMs) using chain-of-thought reasoning often waste substantial compute by producing long, incorrect responses. Abstention can m

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

Model ReleasesDGX agent

arXiv:2603.16120v2 Announce Type: replace Abstract: Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queri

Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions

ApplicationsDGX agent

arXiv:2601.05414v2 Announce Type: replace Abstract: As large language models (LLMs) transition from chat interfaces to integral components of stochastic pipelines and systems approaching general intel

Large Language Models Are Still Misled by Simple Bias Ensembles

Model ReleasesDGX agent

arXiv:2505.16522v3 Announce Type: replace Abstract: With the evolution of large language models (LLMs), their robustness against individual simple biases has been enhanced. However, we observe that th

Latent Abstraction for Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2604.17866v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become a standard approach for enhancing large language models (LLMs) with external knowledge, mitigating hallu

Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering

ResearchDGX agent

arXiv:2604.18567v1 Announce Type: cross Abstract: Large language models frequently commit unrecoverable reasoning errors mid-generation: once a wrong step is taken, subsequent tokens compound the mist

Latent Preference Modeling for Cross-Session Personalized Tool Calling

Model ReleasesDGX agent

arXiv:2604.17886v1 Announce Type: new Abstract: Users often omit essential details in their requests to LLM-based agents, resulting in under-specified inputs for tool use. This poses a fundamental cha

LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations

Model ReleasesDGX agent

arXiv:2509.12539v2 Announce Type: replace-cross Abstract: We present LEAF ('Lightweight Embedding Alignment Framework'), a knowledge distillation framework for text embedding models. A key distinguish

Learning to Control Summaries with Score Ranking

Model ReleasesDGX agent

arXiv:2604.17197v1 Announce Type: new Abstract: Recent advances in summarization research focus on improving summary quality across multiple criteria, such as completeness, conciseness, and faithfulne

Learning to Retrieve User History and Generate User Profiles for Personalized Persuasiveness Prediction

Model ReleasesDGX agent

arXiv:2601.05654v3 Announce Type: replace Abstract: Estimating the persuasiveness of messages is critical in various applications, from recommender systems to safety assessment of LLMs. While it is im

Learning to Seek Help: Dynamic Collaboration Between Small and Large Language Models

Local AiDGX agent

arXiv:2604.17827v1 Announce Type: new Abstract: Large language models (LLMs) offer strong capabilities but raise cost and privacy concerns, whereas small language models (SLMs) facilitate efficient an

Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification

SafetyDGX agent

arXiv:2601.21244v3 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has advanced LLM reasoning, but remains constrained by inefficient exploration under lim

Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection

Model ReleasesDGX agent

arXiv:2506.00955v2 Announce Type: replace Abstract: Sarcasm fundamentally alters meaning through tone and context, yet detecting it in speech remains a challenge due to data scarcity. In addition, exi

LexRel: Benchmarking Legal Relation Extraction for Chinese Civil Cases

Model ReleasesDGX agent

arXiv:2512.12643v2 Announce Type: replace Abstract: Legal relations serve as an important analytical framework for dispute resolution in civil cases. However, legal relations in Chinese civil cases re

LiFT: Does Instruction Fine-Tuning Improve In-Context Learning for Longitudinal Modelling by Large Language Models?

Model ReleasesDGX agent

arXiv:2604.16382v1 Announce Type: new Abstract: Longitudinal NLP tasks require reasoning over temporally ordered text to detect persistence and change in human behavior and opinions. However, in-conte

LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning

Model ReleasesDGX agent

arXiv:2506.00772v2 Announce Type: replace-cross Abstract: Recent studies have shown that supervised fine-tuning of LLMs on a small number of high-quality datasets can yield strong reasoning capabiliti

← Previous
1…107108109110111…129
Next →