AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
4 May 2026

Reasoning-Intensive Regression

Model ReleasesDGX agent

arXiv:2508.21762v3 Announce Type: replace Abstract: AI researchers and practitioners increasingly apply large language models (LLMs) to what we call reasoning-intensive regression (RiR), i.e., deducin

Reinforcement Learning for LLM Post-Training: A Survey

SafetyDGX agent

arXiv:2407.16216v3 Announce Type: replace Abstract: Large language models (LLMs) trained via pretraining and supervised fine-tuning (SFT) can still produce harmful and misaligned outputs, or struggle

ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding, but at What Cost?

SafetyDGX agent

arXiv:2605.00468v1 Announce Type: new Abstract: Plain Language Summaries (PLS) aim to make research accessible to lay readers, but they are typically written in a one-size-fits-all style that ignores


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Representation in large language models

ResearchDGX agent

arXiv:2501.00885v2 Announce Type: replace Abstract: The extraordinary success of recent Large Language Models (LLMs) on a diverse array of tasks has led to an explosion of scientific and philosophical

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

SafetyDGX agent

arXiv:2605.00380v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diver

Rethinking LLM Ensembling from the Perspective of Mixture Models

ResearchDGX agent

arXiv:2605.00419v1 Announce Type: cross Abstract: Model ensembling is a well-established technique for improving the performance of machine learning models. Conventionally, this involves averaging the

Retrieval-Augmented Reasoning for Chartered Accountancy

Model ReleasesDGX agent

arXiv:2605.00257v1 Announce Type: new Abstract: The inception of Large Language Models (LLMs) has catalyzed AI adoption in the finance sector, yet their reliability in complex, jurisdiction-specific t

Reward Modeling from Natural Language Human Feedback

ResearchDGX agent

arXiv:2601.07349v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable reward (RLVR) on preference data has become the mainstream approach for training Generative Reward Models (GR

RouteProfile: Elucidating the Design Space of LLM Profiles for Routing

ResearchDGX agent

arXiv:2605.00180v1 Announce Type: cross Abstract: As the large language model (LLM) ecosystem expands, individual models exhibit varying capabilities across queries, benchmarks, and domains, motivatin

RSAT: Structured Attribution Makes Small Language Models Faithful Table Reasoners

Model ReleasesDGX agent

arXiv:2605.00199v1 Announce Type: new Abstract: When a language model answers a table question, users have no way to verify which cells informed which reasoning steps. We introduce RSAT, a method that

RunAgent: Interpreting Natural-Language Plans with Constraint-Guided Execution

AgentsDGX agent

arXiv:2605.00798v1 Announce Type: cross Abstract: Humans solve problems by executing targeted plans, yet large language models (LLMs) remain unreliable for structured workflow execution. We propose Ru

SC-Taxo: Hierarchical Taxonomy Generation under Semantic Consistency Constraints using Large Language Models

Model ReleasesDGX agent

arXiv:2605.00620v1 Announce Type: new Abstract: Scientific literature is expanding at an unprecedented pace, making it increasingly challenging to efficiently organize and access domain knowledge. A h

Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models

ResearchDGX agent

arXiv:2601.21214v2 Announce Type: replace Abstract: Chain-of-thought (CoT) reasoning has become the standard paradigm for enabling Large Language Models (LLMs) to solve complex problems. However, rece

SCAN: Structured Capability Assessment and Navigation for LLMs

ResearchDGX agent

arXiv:2505.06698v4 Announce Type: replace Abstract: Evaluating Large Language Models (LLMs) has become increasingly important, with automatic evaluation benchmarks gaining prominence as alternatives t

Short Chains, Deep Thoughts: Balancing Reasoning Efficiency and Intra-Segment Capability via Split-Merge Optimization

ResearchDGX agent

arXiv:2602.03141v3 Announce Type: replace Abstract: While Large Reasoning Models (LRMs) have demonstrated impressive capabilities in solving complex tasks through the generation of long reasoning chai

State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning

Model ReleasesDGX agent

arXiv:2605.00206v1 Announce Type: cross Abstract: Current transformers discard their rich latent residual stream between positions, reconstructing latent reasoning context at each new position and lea

Structure-Aware Chunking for Tabular Data in Retrieval-Augmented Generation

ApplicationsDGX agent

arXiv:2605.00318v1 Announce Type: new Abstract: Tabular documents such as CSV and Excel files are widely used in enterprise data pipelines, yet existing chunking strategies for retrieval-augmented gen

Structure Liberates: How Constrained Sensemaking Produces More Novel Research Output

ResearchDGX agent

arXiv:2605.00557v1 Announce Type: new Abstract: Scientific discovery is an extended process of ideation--surveying prior work, forming hypotheses, and refining reasoning--yet existing approaches treat

Structured In-context Environment Scaling for Large Language Model Reasoning

TutorialsDGX agent

arXiv:2509.23330v3 Announce Type: replace Abstract: Large language models (LLMs) have achieved significant advancements in reasoning capabilities through reinforcement learning (RL) via environmental

Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in Dialogue

ApplicationsDGX agent

arXiv:2605.00506v1 Announce Type: new Abstract: We model utterance production as probabilistic cost-sensitive choice over contextual alternatives, using information-theoretic notions of cost. We disti

Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization

ResearchDGX agent

arXiv:2605.00140v1 Announce Type: cross Abstract: We present Activation Residual Hessian Quantization (ARHQ), a post-training weight splitting method designed to mitigate error propagation in low-bit

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

ResearchDGX agent

arXiv:2603.17837v3 Announce Type: replace-cross Abstract: During conversational interactions, humans subconsciously engage in concurrent thinking while listening to a speaker. Although this internal c

Timing is Everything: Temporal Scaffolding of Semantic Surprise in Humor

ResearchDGX agent

arXiv:2605.00143v1 Announce Type: new Abstract: Humor is a fundamental cognitive phenomenon in which humans derive pleasure from the expectation violations and their resolution, exemplifying the brain

Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection

ResearchDGX agent

arXiv:2602.03216v2 Announce Type: replace Abstract: The quadratic complexity of attention remains the central bottleneck in long-context inference for large language models. Prior acceleration methods

ToolGrad: Efficient Tool-use Dataset Generation with Textual 'Gradients'

AgentsDGX agent

arXiv:2508.04086v2 Announce Type: replace Abstract: Prior work synthesizes tool-use LLM datasets by first generating a user query, followed by complex tool-use annotations like depth-first search (DFS

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity

SafetyDGX agent

arXiv:2605.00365v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has achieved substantial gains in single-attempt accuracy (Pass@1) on reasoning tasks, yet often

Unlearning What Matters: Token-Level Attribution for Precise Language Model Unlearning

SafetyDGX agent

arXiv:2605.00364v1 Announce Type: new Abstract: Machine unlearning has emerged as a critical capability for addressing privacy, safety, and regulatory concerns in large language models (LLMs). Existin

VGR: Visual Grounded Reasoning

SafetyDGX agent

arXiv:2506.11991v3 Announce Type: replace-cross Abstract: In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which

ViLegalNLI: Natural Language Inference for Vietnamese Legal Texts

Model ReleasesDGX agent

arXiv:2605.00116v1 Announce Type: new Abstract: In this article, we introduce ViLegalNLI, the first large-scale Vietnamese Natural Language Inference (NLI) dataset specifically constructed for the leg

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback

SafetyDGX agent

arXiv:2605.00155v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has become a core post-training step for aligning large language models, yet the reward signal used

'What Are You Really Trying to Do?': Co-Creating Life Goals from Everyday Computer Use

ResearchDGX agent

arXiv:2605.00497v1 Announce Type: cross Abstract: Recent advances in user modeling make it feasible to conduct open-ended inference over a person's everyday computer use. Despite longstanding visions

What Don't You Understand? Using Large Language Models to Identify and Characterize Student Misconceptions About Challenging Topics

TutorialsDGX agent

arXiv:2605.00294v1 Announce Type: new Abstract: This study presents a systematic approach to identifying and characterizing student misconceptions in online learning environments through a novel combi

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models

Model ReleasesDGX agent

arXiv:2605.00817v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong performance on reasoning benchmarks, but final-answer accuracy alone does not show whether they faithf

When RAG Chatbots Expose Their Backend: An Anonymized Case Study of Privacy and Security Risks in Patient-Facing Medical AI

Model ReleasesDGX agent

arXiv:2605.00796v1 Announce Type: cross Abstract: Background: Patient-facing medical chatbots based on retrieval-augmented generation (RAG) are increasingly promoted to deliver accessible, grounded he

Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions

Model ReleasesDGX agent

arXiv:2605.00226v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly tasked with strategic decision-making under incomplete information, such as in negotiation and policymakin

1 May 2026

A Reproducibility Study of LLM-Based Query Reformulation

Model ReleasesDGX agent

arXiv:2604.27421v1 Announce Type: cross Abstract: Large Language Models (LLMs) are now widely used for query reformulation and expansion in Information Retrieval, with many studies reporting substanti

Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents

SafetyDGX agent

arXiv:2601.01885v2 Announce Type: replace Abstract: Large language model (LLM) agents face fundamental limitations in long-horizon reasoning due to finite context windows, making effective memory mana

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR

Model ReleasesDGX agent

arXiv:2604.27543v1 Announce Type: new Abstract: Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into sh

BatteryPass-12K: The First Dataset for the Novel Digital Battery Passport Conformance Task

Model ReleasesDGX agent

arXiv:2604.26986v1 Announce Type: new Abstract: We introduce a novel task of digital battery passport (DBP) conformance classification and introduce the first public benchmark for the task: BatteryPas

CL-bench Life: Can Language Models Learn from Real-Life Context?

Model ReleasesDGX agent

arXiv:2604.27043v1 Announce Type: new Abstract: Today's AI assistants such as OpenClaw are designed to handle context effectively, making context learning an increasingly important capability for mode

Cross-Lingual Response Consistency in Large Language Models: An ILR-Informed Evaluation of Claude Across Six Languages

Model ReleasesDGX agent

arXiv:2604.27137v1 Announce Type: new Abstract: This paper introduces a systematic evaluation framework grounded in the Interagency Language Roundtable (ILR) Skill Level Descriptions and applies it to

Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability

SafetyDGX agent

arXiv:2602.17469v2 Announce Type: replace Abstract: Recent advances in multilingual representation learning aim to bridge the performance gap between high- and low-resource languages, yet their abilit

Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation

ResearchDGX agent

arXiv:2604.27263v1 Announce Type: new Abstract: Subword tokenization is an essential part of modern large language models (LLMs), yet its specific contributions to training efficiency and model perfor

Do What I Say: A Spoken Prompt Dataset for Instruction-Following

Model ReleasesDGX agent

arXiv:2603.09881v2 Announce Type: replace Abstract: Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompt

DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models

Model ReleasesDGX agent

arXiv:2604.27929v1 Announce Type: new Abstract: With the widespread adoption of large language models (LLMs), understanding their personality representation mechanisms has become critical. As a novel

Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry

SafetyDGX agent

arXiv:2604.27019v1 Announce Type: cross Abstract: Safety-aligned language models must refuse harmful requests without collapsing into broad over-refusal, but the training-time mechanisms behind this t

Ease of dependency distance minimization in star-like structures

ResearchDGX agent

arXiv:2604.28034v1 Announce Type: new Abstract: The syntactic structure of a sentence can be represented as a tree where edges indicate syntactic dependencies between words. When that structure is a s

Emotion-Aware Clickbait Attack in Social Media

ResearchDGX agent

arXiv:2604.27369v1 Announce Type: new Abstract: Clickbait is characterized by disproportionately high emotional intensity relative to informational content, often reinforced by specific structural pat

Entropy of Ukrainian

Model ReleasesDGX agent

arXiv:2604.27534v1 Announce Type: new Abstract: In natural language processing, the entropy of a language is a measure of its unpredictability and complexity. The first study on this subject was condu

EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory

ResearchDGX agent

arXiv:2604.27695v1 Announce Type: cross Abstract: Long-term conversational memory requires retrieving evidence scattered across multiple sessions, yet single-pass retrieval fails on temporal and multi

Exploration Hacking: Can LLMs Learn to Resist RL Training?

SafetyDGX agent

arXiv:2604.28182v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become essential to the post-training of large language models (LLMs) for reasoning, agentic capabilities and alignmen

Exploring Applications of Transfer-State Large Language Models: Cognitive Profiling and Socratic AI Tutoring

SafetyDGX agent

arXiv:2604.27454v1 Announce Type: new Abstract: Large language models (LLMs) sometimes exhibit qualitative shifts in response style under sustained self-referential dialogue conditions (Berg et al., 2

Exploring the Limits of Pruning: Task-Specific Neurons, Model Collapse, and Recovery in Task-Specific Large Language Models

Model ReleasesDGX agent

arXiv:2604.27115v1 Announce Type: new Abstract: Neuron pruning is widely used to reduce the computational cost and parameter footprint of large language models, yet it remains unclear whether neurons

From Coarse to Fine: Benchmarking and Reward Modeling for Writing-Centric Generation Tasks

ResearchDGX agent

arXiv:2604.27453v1 Announce Type: new Abstract: Large language models have achieved remarkable progress in text generation but still struggle with generative writing tasks. In terms of evaluation, exi

From Unstructured to Structured: LLM-Guided Attribute Graphs for Entity Search and Ranking

ApplicationsDGX agent

arXiv:2604.27410v1 Announce Type: cross Abstract: Entity search, i.e., finding the most similar entities to a query entity, faces unique challenges in e-commerce, where product similarity varies acros

Geometry-Calibrated Conformal Abstention for Language Models

ResearchDGX agent

arXiv:2604.27914v1 Announce Type: new Abstract: When language models lack relevant knowledge for a given query, they frequently generate plausible responses that can be hallucinations, rather than adm

HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics

ResearchDGX agent

arXiv:2604.27542v1 Announce Type: new Abstract: Conventionally, Automatic Speech Recognition (ASR) systems are evaluated on their ability to correctly recognize each word contained in a speech signal.

HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats

Model ReleasesDGX agent

arXiv:2604.27470v1 Announce Type: new Abstract: Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited.

Hypencoder Revisited: Reproducibility and Analysis of Non-Linear Scoring for First-Stage Retrieval

ResearchDGX agent

arXiv:2604.27037v1 Announce Type: cross Abstract: The Hypencoder, proposed by Killingback et al., is a retrieval framework that replaces the fixed inner-product scoring function used in standard bi-en

In-context Learning vs. Instruction Tuning: The Case of Small and Multilingual Language Models

SafetyDGX agent

arXiv:2503.01611v3 Announce Type: replace Abstract: Instruction following is a critical ability for Large Language Models to perform downstream tasks. The standard approach to instruction tuning has r

← Previous
1…9192939495…129
Next →