AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

Reasoning-Intensive Regression

DGX agent

arXiv:2508.21762v3 Announce Type: replace Abstract: AI researchers and practitioners increasingly apply large language models (LLMs) to what we call reasoning-intensive regression (RiR), i.e., deducin

model-releasesarxiv-cs-cl
4 May 2026
Safety

Reinforcement Learning for LLM Post-Training: A Survey

DGX agent
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2407.16216v3 Announce Type: replace Abstract: Large language models (LLMs) trained via pretraining and supervised fine-tuning (SFT) can still produce harmful and misaligned outputs, or struggle

safetyarxiv-cs-cl
4 May 2026
Safety

ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding, but at What Cost?

DGX agent

arXiv:2605.00468v1 Announce Type: new Abstract: Plain Language Summaries (PLS) aim to make research accessible to lay readers, but they are typically written in a one-size-fits-all style that ignores

safetyarxiv-cs-cl
4 May 2026
Research

Representation in large language models

DGX agent

arXiv:2501.00885v2 Announce Type: replace Abstract: The extraordinary success of recent Large Language Models (LLMs) on a diverse array of tasks has led to an explosion of scientific and philosophical

researcharxiv-cs-cl
4 May 2026
Safety

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

DGX agent

arXiv:2605.00380v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diver

safetyarxiv-cs-cl
4 May 2026
Research

Rethinking LLM Ensembling from the Perspective of Mixture Models

DGX agent

arXiv:2605.00419v1 Announce Type: cross Abstract: Model ensembling is a well-established technique for improving the performance of machine learning models. Conventionally, this involves averaging the

researcharxiv-cs-cl
4 May 2026
Model Releases

Retrieval-Augmented Reasoning for Chartered Accountancy

DGX agent

arXiv:2605.00257v1 Announce Type: new Abstract: The inception of Large Language Models (LLMs) has catalyzed AI adoption in the finance sector, yet their reliability in complex, jurisdiction-specific t

model-releasesarxiv-cs-cl
4 May 2026
Research

Reward Modeling from Natural Language Human Feedback

DGX agent

arXiv:2601.07349v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable reward (RLVR) on preference data has become the mainstream approach for training Generative Reward Models (GR

researcharxiv-cs-cl
4 May 2026
Research

RouteProfile: Elucidating the Design Space of LLM Profiles for Routing

DGX agent

arXiv:2605.00180v1 Announce Type: cross Abstract: As the large language model (LLM) ecosystem expands, individual models exhibit varying capabilities across queries, benchmarks, and domains, motivatin

researcharxiv-cs-cl
4 May 2026
Model Releases

RSAT: Structured Attribution Makes Small Language Models Faithful Table Reasoners

DGX agent

arXiv:2605.00199v1 Announce Type: new Abstract: When a language model answers a table question, users have no way to verify which cells informed which reasoning steps. We introduce RSAT, a method that

model-releasesarxiv-cs-cl
4 May 2026
Agents

RunAgent: Interpreting Natural-Language Plans with Constraint-Guided Execution

DGX agent

arXiv:2605.00798v1 Announce Type: cross Abstract: Humans solve problems by executing targeted plans, yet large language models (LLMs) remain unreliable for structured workflow execution. We propose Ru

agentsarxiv-cs-cl
4 May 2026
Model Releases

SC-Taxo: Hierarchical Taxonomy Generation under Semantic Consistency Constraints using Large Language Models

DGX agent

arXiv:2605.00620v1 Announce Type: new Abstract: Scientific literature is expanding at an unprecedented pace, making it increasingly challenging to efficiently organize and access domain knowledge. A h

model-releasesarxiv-cs-cl
4 May 2026
Research

Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models

DGX agent

arXiv:2601.21214v2 Announce Type: replace Abstract: Chain-of-thought (CoT) reasoning has become the standard paradigm for enabling Large Language Models (LLMs) to solve complex problems. However, rece

researcharxiv-cs-cl
4 May 2026
Research

SCAN: Structured Capability Assessment and Navigation for LLMs

DGX agent

arXiv:2505.06698v4 Announce Type: replace Abstract: Evaluating Large Language Models (LLMs) has become increasingly important, with automatic evaluation benchmarks gaining prominence as alternatives t

researcharxiv-cs-cl
4 May 2026
Research

Short Chains, Deep Thoughts: Balancing Reasoning Efficiency and Intra-Segment Capability via Split-Merge Optimization

DGX agent

arXiv:2602.03141v3 Announce Type: replace Abstract: While Large Reasoning Models (LRMs) have demonstrated impressive capabilities in solving complex tasks through the generation of long reasoning chai

researcharxiv-cs-cl
4 May 2026
Model Releases

State Stream Transformer (SST) V2: Parallel Training of Nonlinear Recurrence for Latent Space Reasoning

DGX agent

arXiv:2605.00206v1 Announce Type: cross Abstract: Current transformers discard their rich latent residual stream between positions, reconstructing latent reasoning context at each new position and lea

model-releasesarxiv-cs-cl
4 May 2026
Applications

Structure-Aware Chunking for Tabular Data in Retrieval-Augmented Generation

DGX agent

arXiv:2605.00318v1 Announce Type: new Abstract: Tabular documents such as CSV and Excel files are widely used in enterprise data pipelines, yet existing chunking strategies for retrieval-augmented gen

applicationsarxiv-cs-cl
4 May 2026
Research

Structure Liberates: How Constrained Sensemaking Produces More Novel Research Output

DGX agent

arXiv:2605.00557v1 Announce Type: new Abstract: Scientific discovery is an extended process of ideation--surveying prior work, forming hypotheses, and refining reasoning--yet existing approaches treat

researcharxiv-cs-cl
4 May 2026
Tutorials

Structured In-context Environment Scaling for Large Language Model Reasoning

DGX agent

arXiv:2509.23330v3 Announce Type: replace Abstract: Large language models (LLMs) have achieved significant advancements in reasoning capabilities through reinforcement learning (RL) via environmental

tutorialsarxiv-cs-cl
4 May 2026
Applications

Surprisal Minimisation over Goal-directed Alternatives Predicts Production Choice in Dialogue

DGX agent

arXiv:2605.00506v1 Announce Type: new Abstract: We model utterance production as probabilistic cost-sensitive choice over contextual alternatives, using information-theoretic notions of cost. We disti

applicationsarxiv-cs-cl
4 May 2026
Research

Technical Report: Activation Residual Hessian Quantization (ARHQ) for Low-Bit LLM Quantization

DGX agent

arXiv:2605.00140v1 Announce Type: cross Abstract: We present Activation Residual Hessian Quantization (ARHQ), a post-training weight splitting method designed to mitigate error propagation in low-bit

researcharxiv-cs-cl
4 May 2026
Research

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning

DGX agent

arXiv:2603.17837v3 Announce Type: replace-cross Abstract: During conversational interactions, humans subconsciously engage in concurrent thinking while listening to a speaker. Although this internal c

researcharxiv-cs-cl
4 May 2026
Research

Timing is Everything: Temporal Scaffolding of Semantic Surprise in Humor

DGX agent

arXiv:2605.00143v1 Announce Type: new Abstract: Humor is a fundamental cognitive phenomenon in which humans derive pleasure from the expectation violations and their resolution, exemplifying the brain

researcharxiv-cs-cl
4 May 2026
Research

Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection

DGX agent

arXiv:2602.03216v2 Announce Type: replace Abstract: The quadratic complexity of attention remains the central bottleneck in long-context inference for large language models. Prior acceleration methods

researcharxiv-cs-cl
4 May 2026
Agents

ToolGrad: Efficient Tool-use Dataset Generation with Textual 'Gradients'

DGX agent

arXiv:2508.04086v2 Announce Type: replace Abstract: Prior work synthesizes tool-use LLM datasets by first generating a user query, followed by complex tool-use annotations like depth-first search (DFS

agentsarxiv-cs-cl
4 May 2026
Safety

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity

DGX agent

arXiv:2605.00365v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has achieved substantial gains in single-attempt accuracy (Pass@1) on reasoning tasks, yet often

safetyarxiv-cs-cl
4 May 2026
Safety

Unlearning What Matters: Token-Level Attribution for Precise Language Model Unlearning

DGX agent

arXiv:2605.00364v1 Announce Type: new Abstract: Machine unlearning has emerged as a critical capability for addressing privacy, safety, and regulatory concerns in large language models (LLMs). Existin

safetyarxiv-cs-cl
4 May 2026
Safety

VGR: Visual Grounded Reasoning

DGX agent

arXiv:2506.11991v3 Announce Type: replace-cross Abstract: In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which

safetyarxiv-cs-cl
4 May 2026
Model Releases

ViLegalNLI: Natural Language Inference for Vietnamese Legal Texts

DGX agent

arXiv:2605.00116v1 Announce Type: new Abstract: In this article, we introduce ViLegalNLI, the first large-scale Vietnamese Natural Language Inference (NLI) dataset specifically constructed for the leg

model-releasesarxiv-cs-cl
4 May 2026
Safety

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback

DGX agent

arXiv:2605.00155v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) has become a core post-training step for aligning large language models, yet the reward signal used

safetyarxiv-cs-cl
4 May 2026
Research

'What Are You Really Trying to Do?': Co-Creating Life Goals from Everyday Computer Use

DGX agent

arXiv:2605.00497v1 Announce Type: cross Abstract: Recent advances in user modeling make it feasible to conduct open-ended inference over a person's everyday computer use. Despite longstanding visions

researcharxiv-cs-cl
4 May 2026
Tutorials

What Don't You Understand? Using Large Language Models to Identify and Characterize Student Misconceptions About Challenging Topics

DGX agent

arXiv:2605.00294v1 Announce Type: new Abstract: This study presents a systematic approach to identifying and characterizing student misconceptions in online learning environments through a novel combi

tutorialsarxiv-cs-cl
4 May 2026
Model Releases

When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models

DGX agent

arXiv:2605.00817v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong performance on reasoning benchmarks, but final-answer accuracy alone does not show whether they faithf

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

When RAG Chatbots Expose Their Backend: An Anonymized Case Study of Privacy and Security Risks in Patient-Facing Medical AI

DGX agent

arXiv:2605.00796v1 Announce Type: cross Abstract: Background: Patient-facing medical chatbots based on retrieval-augmented generation (RAG) are increasingly promoted to deliver accessible, grounded he

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions

DGX agent

arXiv:2605.00226v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly tasked with strategic decision-making under incomplete information, such as in negotiation and policymakin

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

A Reproducibility Study of LLM-Based Query Reformulation

DGX agent

arXiv:2604.27421v1 Announce Type: cross Abstract: Large Language Models (LLMs) are now widely used for query reformulation and expansion in Information Retrieval, with many studies reporting substanti

model-releasesarxiv-cs-cl
1 May 2026
Safety

Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for Large Language Model Agents

DGX agent

arXiv:2601.01885v2 Announce Type: replace Abstract: Large language model (LLM) agents face fundamental limitations in long-horizon reasoning due to finite context windows, making effective memory mana

safetyarxiv-cs-cl
1 May 2026
Model Releases

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR

DGX agent

arXiv:2604.27543v1 Announce Type: new Abstract: Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into sh

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

BatteryPass-12K: The First Dataset for the Novel Digital Battery Passport Conformance Task

DGX agent

arXiv:2604.26986v1 Announce Type: new Abstract: We introduce a novel task of digital battery passport (DBP) conformance classification and introduce the first public benchmark for the task: BatteryPas

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

CL-bench Life: Can Language Models Learn from Real-Life Context?

DGX agent

arXiv:2604.27043v1 Announce Type: new Abstract: Today's AI assistants such as OpenClaw are designed to handle context effectively, making context learning an increasingly important capability for mode

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

Cross-Lingual Response Consistency in Large Language Models: An ILR-Informed Evaluation of Claude Across Six Languages

DGX agent

arXiv:2604.27137v1 Announce Type: new Abstract: This paper introduces a systematic evaluation framework grounded in the Interagency Language Roundtable (ILR) Skill Level Descriptions and applies it to

model-releasesarxiv-cs-cl
1 May 2026
Safety

Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability

DGX agent

arXiv:2602.17469v2 Announce Type: replace Abstract: Recent advances in multilingual representation learning aim to bridge the performance gap between high- and low-resource languages, yet their abilit

safetyarxiv-cs-cl
1 May 2026
Research

Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation

DGX agent

arXiv:2604.27263v1 Announce Type: new Abstract: Subword tokenization is an essential part of modern large language models (LLMs), yet its specific contributions to training efficiency and model perfor

researcharxiv-cs-cl
1 May 2026
Model Releases

Do What I Say: A Spoken Prompt Dataset for Instruction-Following

DGX agent

arXiv:2603.09881v2 Announce Type: replace Abstract: Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompt

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models

DGX agent

arXiv:2604.27929v1 Announce Type: new Abstract: With the widespread adoption of large language models (LLMs), understanding their personality representation mechanisms has become critical. As a novel

model-releasesarxiv-cs-cl
1 May 2026
Safety

Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry

DGX agent

arXiv:2604.27019v1 Announce Type: cross Abstract: Safety-aligned language models must refuse harmful requests without collapsing into broad over-refusal, but the training-time mechanisms behind this t

safetyarxiv-cs-cl
1 May 2026
Research

Ease of dependency distance minimization in star-like structures

DGX agent

arXiv:2604.28034v1 Announce Type: new Abstract: The syntactic structure of a sentence can be represented as a tree where edges indicate syntactic dependencies between words. When that structure is a s

researcharxiv-cs-cl
1 May 2026
Research

Emotion-Aware Clickbait Attack in Social Media

DGX agent

arXiv:2604.27369v1 Announce Type: new Abstract: Clickbait is characterized by disproportionately high emotional intensity relative to informational content, often reinforced by specific structural pat

researcharxiv-cs-cl
1 May 2026
← Previous
1…114115116117118…161
Next →