AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

Reasoning Gets Harder for LLMs Inside A Dialogue

DGX agent

arXiv:2603.20133v2 Announce Type: replace Abstract: Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that d

model-releasesarxiv-cs-cl
30 Apr 2026
Research

ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2507.21420v3 Announce Type: replace-cross Abstract: The computational cost of training multimodal large language models (MLLMs) grows rapidly with the number of processed tokens. Existing effici

researcharxiv-cs-cl
30 Apr 2026
Research

SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts

DGX agent

arXiv:2604.26506v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly integrated into academic peer review, their vulnerability to adversarial prompts -- adversarial instruc

researcharxiv-cs-cl
30 Apr 2026
Safety

SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling

DGX agent

arXiv:2604.26630v1 Announce Type: new Abstract: Effective mental health counseling is a complex, theory-driven process requiring the simultaneous integration of psychological frameworks, real-time dis

safetyarxiv-cs-cl
30 Apr 2026
Local Ai

Select to Think: Unlocking SLM Potential with Local Sufficiency

DGX agent

arXiv:2604.26940v1 Announce Type: new Abstract: Small language models (SLMs) offer computational efficiency for scalable deployment, yet they often fall short of the reasoning power exhibited by their

local-aiarxiv-cs-cl
30 Apr 2026
Model Releases

Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training

DGX agent

arXiv:2510.20956v2 Announce Type: replace-cross Abstract: We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models (RLMs), which we call self-jailbreaki

model-releasesarxiv-cs-cl
30 Apr 2026
Research

Semantic Embeddings of Chemical Elements for Enhanced Materials Inference and Discovery

DGX agent

arXiv:2502.14912v2 Announce Type: replace Abstract: We present a framework for generating universal semantic embeddings of chemical elements to advance materials inference and discovery. This framewor

researcharxiv-cs-cl
30 Apr 2026
Model Releases

Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens

DGX agent

arXiv:2604.26355v1 Announce Type: new Abstract: Reasoning in Large Language Models incurs significant inference-time compute, yet the token-level information structure of reasoning traces remains unde

model-releasesarxiv-cs-cl
30 Apr 2026
Research

SpecTr-GBV: Multi-Draft Block Verification Accelerating Speculative Decoding

DGX agent

arXiv:2604.25925v1 Announce Type: new Abstract: Autoregressive language models suffer from high inference latency due to their sequential decoding nature. Speculative decoding (SD) mitigates this by e

researcharxiv-cs-cl
30 Apr 2026
Model Releases

StarDrinks: An English and Korean Test Set for SLU Evaluation in a Drink Ordering Scenario

DGX agent

arXiv:2604.26500v1 Announce Type: new Abstract: LLMs and speech assistants are increasingly used for task-oriented interactions, yet their evaluation often relies on controlled scenarios that fail to

model-releasesarxiv-cs-cl
30 Apr 2026
Research

Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images

DGX agent

arXiv:2510.21828v2 Announce Type: replace-cross Abstract: Understanding and reasoning with abstractive information from the visual modality presents significant challenges for current multi-modal larg

researcharxiv-cs-cl
30 Apr 2026
Research

Swap distance minimization shapes the order of subject, object and verb in languages of the world

DGX agent

arXiv:2604.26726v1 Announce Type: new Abstract: Languages of the world vary concerning the order of subject, object and verb. The most frequent dominant orders are SOV and SVO, and researchers have ta

researcharxiv-cs-cl
30 Apr 2026
Model Releases

SWE-Edit: Rethinking Code Editing for Efficient SWE-Agent

DGX agent

arXiv:2604.26102v1 Announce Type: cross Abstract: Large language model agents have achieved remarkable progress on software engineering tasks, yet current approaches suffer from a fundamental context

model-releasesarxiv-cs-cl
30 Apr 2026
Safety

Talent or Luck? Evaluating Attribution Bias in Large Language Models

DGX agent

arXiv:2505.22910v2 Announce Type: replace Abstract: When a student fails an exam, do we tend to blame their effort or the test's difficulty? Attribution, defined as how reasons are assigned to event o

safetyarxiv-cs-cl
30 Apr 2026
Safety

Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment from Heterogeneous Rewards

DGX agent

arXiv:2510.04214v3 Announce Type: replace Abstract: We deploy large language models (LLMs) as business development (BD) agents for persuasive price negotiation in online travel agencies (OTAs). The ag

safetyarxiv-cs-cl
30 Apr 2026
Safety

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation

DGX agent

arXiv:2604.26347v1 Announce Type: cross Abstract: Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring em

safetyarxiv-cs-cl
30 Apr 2026
Model Releases

The Prompt Engineering Report Distilled: Quick Start Guide for Life Sciences

DGX agent

arXiv:2509.11295v2 Announce Type: replace Abstract: Developing effective prompts demands significant cognitive investment to generate reliable, high-quality responses from Large Language Models (LLMs)

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

Theory-Grounded Evaluation Exposes the Authorship Gap in LLM Personalization

DGX agent

arXiv:2604.26460v1 Announce Type: new Abstract: Stylistic personalization - making LLMs write in a specific individual's style, rather than merely adapting to task preferences - lacks evaluation groun

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

Thinking with Drafting: Optical Decompression via Logical Reconstruction

DGX agent

arXiv:2602.11731v2 Announce Type: replace Abstract: Existing multimodal large language models have achieved high-fidelity visual perception and exploratory visual generation. However, a precision para

model-releasesarxiv-cs-cl
30 Apr 2026
Research

Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall

DGX agent

arXiv:2505.13963v3 Announce Type: replace Abstract: Quantization methods are widely used to accelerate inference and streamline the deployment of large language models (LLMs). Although quantization's

researcharxiv-cs-cl
30 Apr 2026
Model Releases

Training-Free Loosely Speculative Decoding: Accepting Semantically Correct Drafts Beyond Exact Match

DGX agent

arXiv:2511.22972v3 Announce Type: replace Abstract: Large language models (LLMs) achieve strong performance across diverse tasks but suffer from high inference latency due to their autoregressive gene

model-releasesarxiv-cs-cl
30 Apr 2026
Safety

Verified Critical Step Optimization for LLM Agents

DGX agent

arXiv:2602.03412v2 Announce Type: replace Abstract: As large language model agents tackle increasingly complex long-horizon tasks, effective post-training becomes critical. Prior work faces fundamenta

safetyarxiv-cs-cl
30 Apr 2026
Model Releases

VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models

DGX agent

arXiv:2505.22897v2 Announce Type: replace Abstract: While bias in large language models (LLMs) is well-studied, similar concerns in vision-language models (VLMs) have received comparatively less atten

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models

DGX agent

arXiv:2510.14438v2 Announce Type: replace Abstract: The hallmark of Deep Research agents lies in compositional reasoning, the capacity to aggregate distributed, heterogeneous information into coherent

model-releasesarxiv-cs-cl
30 Apr 2026
Safety

What Kind of Language is Easy to Language-Model Under Curriculum Learning?

DGX agent

arXiv:2604.26844v1 Announce Type: new Abstract: Many of the thousands of attested languages share common configurations of features, creating a spectrum from typologically very rare (e.g., object-verb

safetyarxiv-cs-cl
30 Apr 2026
Safety

When Annotators Disagree, Topology Explains: Mapper, a Topological Tool for Exploring Text Embedding Geometry and Ambiguity

DGX agent

arXiv:2510.17548v2 Announce Type: replace Abstract: Language models are often evaluated with scalar metrics like accuracy, but such measures fail to capture how models internally represent ambiguity,

safetyarxiv-cs-cl
30 Apr 2026
Research

When Hidden States Drift: Can KV Caches Rescue Long-Range Speculative Decoding?

DGX agent

arXiv:2604.26412v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference, but SOTA hidden-state-based drafters suffer from long-range decay: draft accuracy degrades as the specul

researcharxiv-cs-cl
30 Apr 2026
Research

Zero-Shot to Full-Resource: Cross-lingual Transfer Strategies for Aspect-Based Sentiment Analysis

DGX agent

arXiv:2604.26619v1 Announce Type: new Abstract: Aspect-based Sentiment Analysis (ABSA) extracts fine-grained opinions toward specific aspects within text but remains largely English-focused despite ma

researcharxiv-cs-cl
30 Apr 2026
Safety

A Blueprint for AI-Driven Software Quality: Integrating LLMs with Established Standards

DGX agent

arXiv:2505.13766v5 Announce Type: replace-cross Abstract: Software Quality Assurance (SQA) is critical for delivering reliable, secure, and efficient software products. The Software Quality Assurance

safetyarxiv-cs-cl
29 Apr 2026
Research

A paradox of AI fluency

DGX agent

arXiv:2604.25905v1 Announce Type: new Abstract: How much does a user's skill with AI shape what AI actually delivers for them? This question is critical for users, AI product builders, and society at

researcharxiv-cs-cl
29 Apr 2026
Research

A Survey on LLM-based Conversational User Simulation

DGX agent

arXiv:2604.24977v1 Announce Type: new Abstract: User simulation has long played a vital role in computer science due to its potential to support a wide range of applications. Language, as the primary

researcharxiv-cs-cl
29 Apr 2026
Model Releases

ADE: Adaptive Dictionary Embeddings -- Scaling Multi-Anchor Representations to Large Language Models

DGX agent

arXiv:2604.24940v1 Announce Type: new Abstract: Word embeddings are fundamental to natural language processing, yet traditional approaches represent each word with a single vector, creating representa

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation

DGX agent

arXiv:2602.11224v3 Announce Type: replace-cross Abstract: We present Agent-Diff, a novel benchmarking framework for evaluating agentic Large Language Models (LLMs) on real-world productivity software

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses

DGX agent

arXiv:2604.25850v1 Announce Type: new Abstract: Harnesses have become a central determinant of coding-agent performance, shaping how models interact with repositories, tools, and execution environment

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

An Investigation of Linguistic Biases in LLM-Based Recommendations

DGX agent

arXiv:2604.25456v1 Announce Type: new Abstract: We investigate linguistic biases in LLM-based restaurant and product recommendations given prompts varying across Southern American English (AE), Indian

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

Analyzing LLM Reasoning to Uncover Mental Health Stigma

DGX agent

arXiv:2604.25053v1 Announce Type: new Abstract: While large language models (LLMs) are increasingly being explored for mental health applications, recent studies reveal that they can exhibit stigma to

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering

DGX agent

arXiv:2601.12248v2 Announce Type: replace-cross Abstract: Recent advances in audio-aware large language models have shown strong performance on audio question answering. However, existing benchmarks m

model-releasesarxiv-cs-cl
29 Apr 2026
Research

Backtranslation Augmented Direct Preference Optimization for Neural Machine Translation

DGX agent

arXiv:2604.25702v1 Announce Type: new Abstract: Contemporary neural machine translation (NMT) systems are almost exclusively built by training on supervised parallel data. Despite the tremendous progr

researcharxiv-cs-cl
29 Apr 2026
Safety

BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate

DGX agent

arXiv:2604.25203v1 Announce Type: new Abstract: Deploying guardrails for custom policies remains challenging, as generic safety models fail to capture task-specific requirements, while prompting LLMs

safetyarxiv-cs-cl
29 Apr 2026
Tutorials

Barriers to Universal Reasoning With Transformers (And How to Overcome Them)

DGX agent

arXiv:2604.25800v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) has been shown to empirically improve Transformers' performance, and theoretically increase their expressivity to Turing comple

tutorialsarxiv-cs-cl
29 Apr 2026
Model Releases

Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance

DGX agent

arXiv:2604.25249v1 Announce Type: new Abstract: Detecting sandbagging--the deliberate underperformance on capability evaluations--is an open problem in AI safety. We tested whether symptom validity te

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks

DGX agent

arXiv:2604.24955v1 Announce Type: new Abstract: As benchmarks grow in complexity, many apparent agent failures are not failures of the agent at all - they are failures of the benchmark itself: broken

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

Benchmarking and Adapting On-Device LLMs for Clinical Decision Support

DGX agent

arXiv:2601.03266v2 Announce Type: replace Abstract: Large language models (LLMs) have rapidly advanced in clinical decision-making, yet the deployment of proprietary systems is hindered by privacy con

model-releasesarxiv-cs-cl
29 Apr 2026
Research

Benchmarking Logistic Regression, SVM, and LightGBM Against BiLSTM with Attention for Sentiment Analysis on Indonesian Product Reviews

DGX agent

arXiv:2604.25452v1 Announce Type: new Abstract: Sentiment analysis of product reviews on e-commerce platforms plays a critical role in automatically understanding customer satisfaction and providing a

researcharxiv-cs-cl
29 Apr 2026
Research

Benchmarking PyCaret AutoML Against IndoBERT Fine-Tuning for Sentiment Analysis on Indonesian IKN Twitter Data

DGX agent

arXiv:2604.25392v1 Announce Type: new Abstract: This paper benchmarks a classical machine learning approach based on PyCaret AutoML against a deep learning approach based on IndoBERT fine-tuning for b

researcharxiv-cs-cl
29 Apr 2026
Model Releases

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal

DGX agent

arXiv:2509.09708v3 Announce Type: replace Abstract: Refusal on harmful prompts is a key safety behaviour in instruction-tuned large language models (LLMs), yet the internal causes of this behaviour re

model-releasesarxiv-cs-cl
29 Apr 2026
Model Releases

BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding

DGX agent

arXiv:2512.12087v3 Announce Type: replace Abstract: The growing demand for long-context inference capabilities in Large Language Models (LLMs) has intensified the computational and memory bottlenecks

model-releasesarxiv-cs-cl
29 Apr 2026
Research

Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation

DGX agent

arXiv:2604.25580v1 Announce Type: new Abstract: The closure of Perspective API at the end of 2026 discards what has functioned as the de facto standard for automated toxicity measurement in NLP, CSS,

researcharxiv-cs-cl
29 Apr 2026
← Previous
1…117118119120121…161
Next →