AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Local Ai

Characterizing the Expressivity of Local Attention in Transformers

DGX agent

arXiv:2605.00768v1 Announce Type: new Abstract: The transformer is the most popular neural architecture for language modeling. The cornerstone of the transformer is its global attention mechanism, whi

local-aiarxiv-cs-cl
4 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments

DGX agent

arXiv:2505.09901v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate or automate human behavior in complex sequential decision-making settings. A na

researcharxiv-cs-cl
4 May 2026
Research

Confidence Estimation in Automatic Short Answer Grading with LLMs

DGX agent

arXiv:2605.00200v1 Announce Type: new Abstract: Automatic Short Answer Grading (ASAG) with generative large language models (LLMs) has recently demonstrated strong performance without task-specific fi

researcharxiv-cs-cl
4 May 2026
Model Releases

ControBench: An Interaction-Aware Benchmark for Controversial Discourse Analysis on Social Networks

DGX agent

arXiv:2605.00513v1 Announce Type: new Abstract: Understanding how people argue across ideological divides online is important for studying political polarization, misinformation, and content moderatio

model-releasesarxiv-cs-cl
4 May 2026
Research

Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues

DGX agent

arXiv:2605.00119v1 Announce Type: new Abstract: There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. M

researcharxiv-cs-cl
4 May 2026
Research

Directed Social Regard: Surfacing Targeted Advocacy, Opposition, Aid, Harms, and Victimization in Online Media

DGX agent

arXiv:2605.00776v1 Announce Type: new Abstract: The language in online platforms, influence operations, and political rhetoric frequently directs a mix of pro-social sentiment (e.g., advocacy, helpful

researcharxiv-cs-cl
4 May 2026
Safety

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment

DGX agent

arXiv:2506.00166v2 Announce Type: replace-cross Abstract: Existing paradigms for ensuring AI safety, such as guardrail models and alignment training, often compromise either inference efficiency or de

safetyarxiv-cs-cl
4 May 2026
Local Ai

EGREFINE: An Execution-Grounded Optimization Framework for Text-to-SQL Schema Refinement

DGX agent

arXiv:2605.00628v1 Announce Type: cross Abstract: Text-to-SQL enables non-expert users to query databases in natural language, yet real-world schemas often suffer from ambiguous, abbreviated, or incon

local-aiarxiv-cs-cl
4 May 2026
Research

Escaping Mode Collapse in LLM Generation via Geometric Regulation

DGX agent

arXiv:2605.00435v1 Announce Type: new Abstract: Mode collapse is a persistent challenge in generative modeling and appears in autoregressive text generation as behaviors ranging from explicit looping

researcharxiv-cs-cl
4 May 2026
Safety

Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory

DGX agent

arXiv:2605.00238v1 Announce Type: new Abstract: Automated short answer grading (ASAG) with large language models (LLMs) is commonly evaluated with aggregate metrics such as macro-F1 and Cohen's kappa.

safetyarxiv-cs-cl
4 May 2026
Applications

Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics

DGX agent

arXiv:2512.01020v2 Announce Type: replace-cross Abstract: Evaluating the quality of LLM-generated reasoning traces in expert domains (e.g., law) is essential for ensuring credibility and explainabilit

applicationsarxiv-cs-cl
4 May 2026
Model Releases

ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation

DGX agent

arXiv:2507.14201v3 Announce Type: replace-cross Abstract: We present ExCyTIn-Bench, the first benchmark to Evaluate an LLM agent X on the task of Cyber Threat Investigation through security questions

model-releasesarxiv-cs-cl
4 May 2026
Safety

Exploring LLM biases to manipulate AI search overview

DGX agent

arXiv:2605.00012v1 Announce Type: cross Abstract: Modern large language models (LLMs) are used in many business applications in general, and specifically in web search systems and applications that ge

safetyarxiv-cs-cl
4 May 2026
Model Releases

Exploring the System 1 Thinking Capability of Large Reasoning Models

DGX agent

arXiv:2504.10368v4 Announce Type: replace Abstract: This paper explores the system 1 thinking capability of Large Reasoning Models (LRMs), the intuitive ability to respond efficiently with minimal tok

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios

DGX agent

arXiv:2605.00706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

FollowTable: A Benchmark for Instruction-Following Table Retrieval

DGX agent

arXiv:2605.00400v1 Announce Type: cross Abstract: Table Retrieval (TR) has traditionally been formulated as an ad-hoc retrieval problem, where relevance is primarily determined by topical semantic sim

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter Editing

DGX agent

arXiv:2605.00358v1 Announce Type: new Abstract: LLM parameter editing methods commonly rely on computing an ideal target hidden-state at a target layer (referred as anchor point) and distributing the

model-releasesarxiv-cs-cl
4 May 2026
Research

H-RAG at SemEval-2026 Task 8: Hierarchical Parent-Child Retrieval for Multi-Turn RAG Conversations

DGX agent

arXiv:2605.00631v1 Announce Type: new Abstract: We present H-RAG, our submission to SemEval-2026 Task 8 (MTRAGEval), addressing both Task A (Retrieval) and Task C (Generation with Retrieved Passages).

researcharxiv-cs-cl
4 May 2026
Model Releases

How Frontier LLMs Adapt to Neurodivergence Context: A Measurement Framework for Surface vs. Structural Change in System-Prompted Responses

DGX agent

arXiv:2605.00113v1 Announce Type: new Abstract: We examine if frontier chat-based large language models (LLMs) adjust their outputs based on neurodivergence (ND) context in system prompts and describe

model-releasesarxiv-cs-cl
4 May 2026
Safety

How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework

DGX agent

arXiv:2605.00269v1 Announce Type: new Abstract: Recent white-box OOD detection methods for LLMs -- including CED, RAUQ, and WildGuard confidence scores -- appear effective, but we show they are struct

safetyarxiv-cs-cl
4 May 2026
Safety

Impact of Task Phrasing on Presumptions in Large Language Models

DGX agent

arXiv:2605.00436v1 Announce Type: new Abstract: Concerns with the safety and reliability of applying large-language models (LLMs) in unpredictable real-world applications motivate this study, which ex

safetyarxiv-cs-cl
4 May 2026
Model Releases

InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information

DGX agent

arXiv:2508.07630v2 Announce Type: replace Abstract: We introduce InterChart, a diagnostic benchmark that evaluates how well vision-language models (VLMs) reason across multiple related charts, a task

model-releasesarxiv-cs-cl
4 May 2026
Research

Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus

DGX agent

arXiv:2605.00618v1 Announce Type: new Abstract: We investigate the extent to which cosine similarity between paragraph embeddings is invariant under machine translation, using the Manifesto Corpus of

researcharxiv-cs-cl
4 May 2026
Model Releases

Knowing When to Defer: Selective Prediction for Responsible Knowledge Tracing

DGX agent

arXiv:2509.21514v3 Announce Type: replace-cross Abstract: Research on Knowledge Tracing (KT) models traditionally focuses on improving predictive accuracy. However, responsible real-world deployment r

model-releasesarxiv-cs-cl
4 May 2026
Applications

Language-free Experience at Expo 2025 Osaka

DGX agent

arXiv:2605.00373v1 Announce Type: new Abstract: In line with the Global Communication Plan 2025, we have pursued the development of multilingual translation technologies to realize a language-barrier-

applicationsarxiv-cs-cl
4 May 2026
Applications

Language Models Struggle to Use Representations Learned In-Context

DGX agent

arXiv:2602.04212v2 Announce Type: replace Abstract: Though large language models (LLMs) have enabled great success across a wide variety of tasks, they still appear to fall short of one of the loftier

applicationsarxiv-cs-cl
4 May 2026
Research

LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation

DGX agent

arXiv:2605.00777v1 Announce Type: cross Abstract: A speaker encoder used in multilingual voice cloning should treat the same speaker identically regardless of which script the audio was uttered in. Of

researcharxiv-cs-cl
4 May 2026
Model Releases

Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation

DGX agent

arXiv:2510.19897v2 Announce Type: replace Abstract: We investigate how agents built on pretrained large language models (LLMs) can learn target classification functions from labeled examples without p

model-releasesarxiv-cs-cl
4 May 2026
Safety

Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory

DGX agent

arXiv:2605.00702v1 Announce Type: new Abstract: Large language model (LLM) agents require long-term user memory for consistent personalization, but limited context windows hinder tracking evolving pre

safetyarxiv-cs-cl
4 May 2026
Model Releases

Lightweight Domain Adaptation of a Large Language Model for Legal Assistance in the Indian Context

DGX agent

arXiv:2505.22003v2 Announce Type: replace Abstract: In India, access to legal assistance for the general public has been observed to have a critical gap, as many citizens are not able to take full adv

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

LLM-Oriented Information Retrieval: A Denoising-First Perspective

DGX agent

arXiv:2605.00505v1 Announce Type: cross Abstract: Modern information retrieval (IR) is no longer consumed primarily by humans but increasingly by large language models (LLMs) via retrieval-augmented g

model-releasesarxiv-cs-cl
4 May 2026
Research

Lost in State Space: Probing Frozen Mamba Representations

DGX agent

arXiv:2605.00253v1 Announce Type: new Abstract: Mamba's recurrent state h_t is, by construction, a compressed summary of every token seen so far. This raises a tempting hypothesis: if we extract token

researcharxiv-cs-cl
4 May 2026
Research

Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding

DGX agent

arXiv:2605.00342v1 Announce Type: new Abstract: Tree-based speculative decoding accelerates autoregressive generation by verifying multiple draft candidates in parallel, but this advantage weakens for

researcharxiv-cs-cl
4 May 2026
Agents

Memory in the LLM Era: Modular Architectures and Strategies in a Unified Framework

DGX agent

arXiv:2604.01707v2 Announce Type: replace Abstract: Memory emerges as the core module in the large language model (LLM)-based agents for long-horizon complex tasks (e.g., multi-turn dialogue, game pla

agentsarxiv-cs-cl
4 May 2026
Safety

MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents

DGX agent

arXiv:2605.00356v1 Announce Type: new Abstract: Long-term conversational agents must decide which turns to store in external memory, yet recent systems rely on autoregressive LLM generation at every t

safetyarxiv-cs-cl
4 May 2026
Model Releases

ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering

DGX agent

arXiv:2505.23723v2 Announce Type: replace Abstract: The emergence of large language model (LLM)-based agents has significantly advanced the development of autonomous machine learning (ML) engineering.

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models

DGX agent

arXiv:2605.00689v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed in cross-linguistic contexts, ensuring safety in diverse regulatory and cultural environments

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

MoDAl: Self-Supervised Neural Modality Discovery via Decorrelation for Speech Neuroprosthesis

DGX agent

arXiv:2605.00025v1 Announce Type: cross Abstract: Speech neuroprosthesis systems decode intended speech from neural activity in the absence of audible output, offering a path to restoring communicatio

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus

DGX agent

arXiv:2605.00086v1 Announce Type: new Abstract: High-quality corpora are essential for advancing Natural Language Processing (NLP) in Portuguese. Building on previous encoder-only models such as BERTi

model-releasesarxiv-cs-cl
4 May 2026
Research

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning

DGX agent

arXiv:2605.00347v1 Announce Type: cross Abstract: Given the rapidly growing capabilities of vision-language models (VLMs), extending them to interactive decision-making tasks such as video games has e

researcharxiv-cs-cl
4 May 2026
Agents

On the Role of Artificial Intelligence in Human-Machine Symbiosis

DGX agent

arXiv:2605.00440v1 Announce Type: cross Abstract: The evolution of artificial intelligence (AI) has rendered the boundary between humanity and computational machinery increasingly ambiguous. In the pr

agentsarxiv-cs-cl
4 May 2026
Model Releases

Peek2: Regex-free Byte-level Byte-Pair Encoding Pretokenizer for LLM Inference on Edge Devices

DGX agent

arXiv:2601.05833v2 Announce Type: replace Abstract: Pretokenization is a crucial, sequential pass in Byte-level BPE tokenizers, yet little work has been done to optimize it for edge-side inference. Ou

model-releasesarxiv-cs-cl
4 May 2026
Safety

Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations

DGX agent

arXiv:2605.00227v1 Announce Type: new Abstract: There are growing concerns about the risks posed by AI companion applications designed for emotional engagement. Existing safety evaluations often rely

safetyarxiv-cs-cl
4 May 2026
Safety

PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning

DGX agent

arXiv:2510.26020v2 Announce Type: replace Abstract: Multi-tool-integrated reasoning enables LLM-empowered tool-use agents to solve complex tasks by interleaving natural-language reasoning with calls t

safetyarxiv-cs-cl
4 May 2026
Model Releases

Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation

DGX agent

arXiv:2601.06600v2 Announce Type: replace Abstract: Short-video platforms have become major channels for misinformation, where deceptive claims frequently leverage visual experiments and social cues.

model-releasesarxiv-cs-cl
4 May 2026
Safety

Prompt-Induced Score Variance in Zero-Shot Binary Vision-Language Safety Classification

DGX agent

arXiv:2605.00326v1 Announce Type: new Abstract: Single-prompt first-token probabilities from zero-shot vision-language model (VLM) safety classifiers are treated as decision scores, but we show they a

safetyarxiv-cs-cl
4 May 2026
Model Releases

Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment

DGX agent

arXiv:2605.00022v1 Announce Type: new Abstract: The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly. To fill

model-releasesarxiv-cs-cl
4 May 2026
Local Ai

RadLite: Multi-Task LoRA Fine-Tuning of Small Language Models for CPU-Deployable Radiology AI

DGX agent

arXiv:2605.00421v1 Announce Type: new Abstract: Large language models (LLMs) show promise in radiology but their deployment is limited by computational requirements that preclude use in resource-const

local-aiarxiv-cs-cl
4 May 2026
← Previous
1…113114115116117…161
Next →