AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
4 May 2026

Are You the A-hole? A Fair, Multi-Perspective Ethical Reasoning Framework

ApplicationsDGX agent

arXiv:2605.00270v1 Announce Type: new Abstract: Standard methods for aggregating natural language judgments, such as majority voting, often fail to produce logically consistent results when applied to

BanglaSocialBench: A Benchmark for Evaluating Sociopragmatic and Cultural Alignment of LLMs in Bangladeshi Social Interaction

Model ReleasesDGX agent

arXiv:2603.15949v3 Announce Type: replace Abstract: Large Language Models have demonstrated strong multilingual fluency, yet fluency alone does not guarantee socially appropriate language use. In high

Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.00674v1 Announce Type: new Abstract: Large language models (LLMs) are becoming increasingly capable mathematical collaborators, but static benchmarks are no longer sufficient for evaluating

Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe

ResearchDGX agent

arXiv:2605.00607v1 Announce Type: new Abstract: Probing is widely used to study which features can be decoded from language model representations. However, the common decoding probe approach has two l

Bias in Large Language Models: Origin, Evaluation, and Mitigation

SafetyDGX agent

arXiv:2411.10915v2 Announce Type: replace Abstract: Large Language Models (LLMs) have revolutionized natural language processing, but their susceptibility to biases poses significant challenges. This

Block-wise Codeword Embedding for Reliable Multi-bit Text Watermarking

Local AiDGX agent

arXiv:2605.00348v1 Announce Type: cross Abstract: Recent multi-bit watermarking methods for large language models (LLMs) prioritize capacity over reliability, often conflating decoding with detection.

Borrowed Geometry: Computational Reuse of Frozen Text-Pretrained Transformer Weights Across Modalities

Model ReleasesDGX agent

arXiv:2605.00333v1 Announce Type: cross Abstract: Frozen Gemma 4 31B weights pretrained exclusively on text tokens, unmodified, transfer across modality boundaries through a thin trainable interface.

Bring Your Own Prompts: Use-Case-Specific Bias and Fairness Evaluation for LLMs

Model ReleasesDGX agent

arXiv:2407.10853v5 Announce Type: replace Abstract: Bias and fairness risks in Large Language Models (LLMs) vary substantially across deployment contexts, yet existing approaches lack systematic guida

Budget-Aware Routing for Long Clinical Text

ResearchDGX agent

arXiv:2605.00336v1 Announce Type: new Abstract: A key challenge for large language models is token cost per query and overall deployment cost. Clinical inputs are long, heterogeneous, and often redund

Build, Judge, Optimize: A Blueprint for Continuous Improvement of Multi-Agent Consumer Assistants

AgentsDGX agent

arXiv:2603.03565v2 Announce Type: replace-cross Abstract: Conversational shopping assistants (CSAs) represent a compelling application of agentic AI, but moving from prototype to production reveals tw

Can Coding Agents Reproduce Findings in Computational Materials Science?

Model ReleasesDGX agent

arXiv:2605.00803v1 Announce Type: cross Abstract: Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering be

Can Small Language Models Handle Context-Summarized Multi-Turn Customer-Service QA? A Synthetic Data-Driven Comparative Evaluation

SafetyDGX agent

arXiv:2602.00665v3 Announce Type: replace Abstract: Customer-service question answering (QA) systems increasingly rely on conversational language understanding. While Large Language Models (LLMs) achi

Characterizing the Expressivity of Local Attention in Transformers

Local AiDGX agent

arXiv:2605.00768v1 Announce Type: new Abstract: The transformer is the most popular neural architecture for language modeling. The cornerstone of the transformer is its global attention mechanism, whi

Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments

ResearchDGX agent

arXiv:2505.09901v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to simulate or automate human behavior in complex sequential decision-making settings. A na

Confidence Estimation in Automatic Short Answer Grading with LLMs

ResearchDGX agent

arXiv:2605.00200v1 Announce Type: new Abstract: Automatic Short Answer Grading (ASAG) with generative large language models (LLMs) has recently demonstrated strong performance without task-specific fi

ControBench: An Interaction-Aware Benchmark for Controversial Discourse Analysis on Social Networks

Model ReleasesDGX agent

arXiv:2605.00513v1 Announce Type: new Abstract: Understanding how people argue across ideological divides online is important for studying political polarization, misinformation, and content moderatio

Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues

ResearchDGX agent

arXiv:2605.00119v1 Announce Type: new Abstract: There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. M

Directed Social Regard: Surfacing Targeted Advocacy, Opposition, Aid, Harms, and Victimization in Online Media

ResearchDGX agent

arXiv:2605.00776v1 Announce Type: new Abstract: The language in online platforms, influence operations, and political rhetoric frequently directs a mix of pro-social sentiment (e.g., advocacy, helpful

Disentangled Safety Adapters Enable Efficient Guardrails and Flexible Inference-Time Alignment

SafetyDGX agent

arXiv:2506.00166v2 Announce Type: replace-cross Abstract: Existing paradigms for ensuring AI safety, such as guardrail models and alignment training, often compromise either inference efficiency or de

EGREFINE: An Execution-Grounded Optimization Framework for Text-to-SQL Schema Refinement

Local AiDGX agent

arXiv:2605.00628v1 Announce Type: cross Abstract: Text-to-SQL enables non-expert users to query databases in natural language, yet real-world schemas often suffer from ambiguous, abbreviated, or incon

Escaping Mode Collapse in LLM Generation via Geometric Regulation

ResearchDGX agent

arXiv:2605.00435v1 Announce Type: new Abstract: Mode collapse is a persistent challenge in generative modeling and appears in autoregressive text generation as behaviors ranging from explicit looping

Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory

SafetyDGX agent

arXiv:2605.00238v1 Announce Type: new Abstract: Automated short answer grading (ASAG) with large language models (LLMs) is commonly evaluated with aggregate metrics such as macro-F1 and Cohen's kappa.

Evaluating Legal Reasoning Traces with Legal Issue Tree Rubrics

ApplicationsDGX agent

arXiv:2512.01020v2 Announce Type: replace-cross Abstract: Evaluating the quality of LLM-generated reasoning traces in expert domains (e.g., law) is essential for ensuring credibility and explainabilit

ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation

Model ReleasesDGX agent

arXiv:2507.14201v3 Announce Type: replace-cross Abstract: We present ExCyTIn-Bench, the first benchmark to Evaluate an LLM agent X on the task of Cyber Threat Investigation through security questions

Exploring LLM biases to manipulate AI search overview

SafetyDGX agent

arXiv:2605.00012v1 Announce Type: cross Abstract: Modern large language models (LLMs) are used in many business applications in general, and specifically in web search systems and applications that ge

Exploring the System 1 Thinking Capability of Large Reasoning Models

Model ReleasesDGX agent

arXiv:2504.10368v4 Announce Type: replace Abstract: This paper explores the system 1 thinking capability of Large Reasoning Models (LRMs), the intuitive ability to respond efficiently with minimal tok

FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios

Model ReleasesDGX agent

arXiv:2605.00706v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal

FollowTable: A Benchmark for Instruction-Following Table Retrieval

Model ReleasesDGX agent

arXiv:2605.00400v1 Announce Type: cross Abstract: Table Retrieval (TR) has traditionally been formulated as an ad-hoc retrieval problem, where relevance is primarily determined by topical semantic sim

From Backward Spreading to Forward Replay: Revisiting Target Construction in LLM Parameter Editing

Model ReleasesDGX agent

arXiv:2605.00358v1 Announce Type: new Abstract: LLM parameter editing methods commonly rely on computing an ideal target hidden-state at a target layer (referred as anchor point) and distributing the

H-RAG at SemEval-2026 Task 8: Hierarchical Parent-Child Retrieval for Multi-Turn RAG Conversations

ResearchDGX agent

arXiv:2605.00631v1 Announce Type: new Abstract: We present H-RAG, our submission to SemEval-2026 Task 8 (MTRAGEval), addressing both Task A (Retrieval) and Task C (Generation with Retrieved Passages).

How Frontier LLMs Adapt to Neurodivergence Context: A Measurement Framework for Surface vs. Structural Change in System-Prompted Responses

Model ReleasesDGX agent

arXiv:2605.00113v1 Announce Type: new Abstract: We examine if frontier chat-based large language models (LLMs) adjust their outputs based on neurodivergence (ND) context in system prompts and describe

How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework

SafetyDGX agent

arXiv:2605.00269v1 Announce Type: new Abstract: Recent white-box OOD detection methods for LLMs -- including CED, RAUQ, and WildGuard confidence scores -- appear effective, but we show they are struct

Impact of Task Phrasing on Presumptions in Large Language Models

SafetyDGX agent

arXiv:2605.00436v1 Announce Type: new Abstract: Concerns with the safety and reliability of applying large-language models (LLMs) in unpredictable real-world applications motivate this study, which ex

InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information

Model ReleasesDGX agent

arXiv:2508.07630v2 Announce Type: replace Abstract: We introduce InterChart, a diagnostic benchmark that evaluates how well vision-language models (VLMs) reason across multiple related charts, a task

Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus

ResearchDGX agent

arXiv:2605.00618v1 Announce Type: new Abstract: We investigate the extent to which cosine similarity between paragraph embeddings is invariant under machine translation, using the Manifesto Corpus of

Knowing When to Defer: Selective Prediction for Responsible Knowledge Tracing

Model ReleasesDGX agent

arXiv:2509.21514v3 Announce Type: replace-cross Abstract: Research on Knowledge Tracing (KT) models traditionally focuses on improving predictive accuracy. However, responsible real-world deployment r

Language-free Experience at Expo 2025 Osaka

ApplicationsDGX agent

arXiv:2605.00373v1 Announce Type: new Abstract: In line with the Global Communication Plan 2025, we have pursued the development of multilingual translation technologies to realize a language-barrier-

Language Models Struggle to Use Representations Learned In-Context

ApplicationsDGX agent

arXiv:2602.04212v2 Announce Type: replace Abstract: Though large language models (LLMs) have enabled great success across a wide variety of tasks, they still appear to fall short of one of the loftier

LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservation

ResearchDGX agent

arXiv:2605.00777v1 Announce Type: cross Abstract: A speaker encoder used in multilingual voice cloning should treat the same speaker identically regardless of which script the audio was uttered in. Of

Learning from Supervision with Semantic and Episodic Memory: A Reflective Approach to Agent Adaptation

Model ReleasesDGX agent

arXiv:2510.19897v2 Announce Type: replace Abstract: We investigate how agents built on pretrained large language models (LLMs) can learn target classification functions from labeled examples without p

Learning How and What to Memorize: Cognition-Inspired Two-Stage Optimization for Evolving Memory

SafetyDGX agent

arXiv:2605.00702v1 Announce Type: new Abstract: Large language model (LLM) agents require long-term user memory for consistent personalization, but limited context windows hinder tracking evolving pre

Lightweight Domain Adaptation of a Large Language Model for Legal Assistance in the Indian Context

Model ReleasesDGX agent

arXiv:2505.22003v2 Announce Type: replace Abstract: In India, access to legal assistance for the general public has been observed to have a critical gap, as many citizens are not able to take full adv

LLM-Oriented Information Retrieval: A Denoising-First Perspective

Model ReleasesDGX agent

arXiv:2605.00505v1 Announce Type: cross Abstract: Modern information retrieval (IR) is no longer consumed primarily by humans but increasingly by large language models (LLMs) via retrieval-augmented g

Lost in State Space: Probing Frozen Mamba Representations

ResearchDGX agent

arXiv:2605.00253v1 Announce Type: new Abstract: Mamba's recurrent state h_t is, by construction, a compressed summary of every token seen so far. This raises a tempting hypothesis: if we extract token

Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding

ResearchDGX agent

arXiv:2605.00342v1 Announce Type: new Abstract: Tree-based speculative decoding accelerates autoregressive generation by verifying multiple draft candidates in parallel, but this advantage weakens for

Memory in the LLM Era: Modular Architectures and Strategies in a Unified Framework

AgentsDGX agent

arXiv:2604.01707v2 Announce Type: replace Abstract: Memory emerges as the core module in the large language model (LLM)-based agents for long-horizon complex tasks (e.g., multi-turn dialogue, game pla

MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents

SafetyDGX agent

arXiv:2605.00356v1 Announce Type: new Abstract: Long-term conversational agents must decide which turns to store in external memory, yet recent systems rely on autoregressive LLM generation at every t

ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering

Model ReleasesDGX agent

arXiv:2505.23723v2 Announce Type: replace Abstract: The emergence of large language model (LLM)-based agents has significantly advanced the development of autonomous machine learning (ML) engineering.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models

Model ReleasesDGX agent

arXiv:2605.00689v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed in cross-linguistic contexts, ensuring safety in diverse regulatory and cultural environments

MoDAl: Self-Supervised Neural Modality Discovery via Decorrelation for Speech Neuroprosthesis

Model ReleasesDGX agent

arXiv:2605.00025v1 Announce Type: cross Abstract: Speech neuroprosthesis systems decode intended speech from neural activity in the absence of audible output, offering a path to restoring communicatio

NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus

Model ReleasesDGX agent

arXiv:2605.00086v1 Announce Type: new Abstract: High-quality corpora are essential for advancing Natural Language Processing (NLP) in Portuguese. Building on previous encoder-only models such as BERTi

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning

ResearchDGX agent

arXiv:2605.00347v1 Announce Type: cross Abstract: Given the rapidly growing capabilities of vision-language models (VLMs), extending them to interactive decision-making tasks such as video games has e

On the Role of Artificial Intelligence in Human-Machine Symbiosis

AgentsDGX agent

arXiv:2605.00440v1 Announce Type: cross Abstract: The evolution of artificial intelligence (AI) has rendered the boundary between humanity and computational machinery increasingly ambiguous. In the pr

Peek2: Regex-free Byte-level Byte-Pair Encoding Pretokenizer for LLM Inference on Edge Devices

Model ReleasesDGX agent

arXiv:2601.05833v2 Announce Type: replace Abstract: Pretokenization is a crucial, sequential pass in Byte-level BPE tokenizers, yet little work has been done to optimize it for edge-side inference. Ou

Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations

SafetyDGX agent

arXiv:2605.00227v1 Announce Type: new Abstract: There are growing concerns about the risks posed by AI companion applications designed for emotional engagement. Existing safety evaluations often rely

PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning

SafetyDGX agent

arXiv:2510.26020v2 Announce Type: replace Abstract: Multi-tool-integrated reasoning enables LLM-empowered tool-use agents to solve complex tasks by interleaving natural-language reasoning with calls t

Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation

Model ReleasesDGX agent

arXiv:2601.06600v2 Announce Type: replace Abstract: Short-video platforms have become major channels for misinformation, where deceptive claims frequently leverage visual experiments and social cues.

Prompt-Induced Score Variance in Zero-Shot Binary Vision-Language Safety Classification

SafetyDGX agent

arXiv:2605.00326v1 Announce Type: new Abstract: Single-prompt first-token probabilities from zero-shot vision-language model (VLM) safety classifiers are treated as decision scores, but we show they a

Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment

Model ReleasesDGX agent

arXiv:2605.00022v1 Announce Type: new Abstract: The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly. To fill

RadLite: Multi-Task LoRA Fine-Tuning of Small Language Models for CPU-Deployable Radiology AI

Local AiDGX agent

arXiv:2605.00421v1 Announce Type: new Abstract: Large language models (LLMs) show promise in radiology but their deployment is limited by computational requirements that preclude use in resource-const

← Previous
1…9091929394…129
Next →