AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
19 May 2026

Provable Knowledge Acquisition and Extraction in One-Layer Transformers

Model ReleasesDGX agent

arXiv:2508.00901v4 Announce Type: replace-cross Abstract: Large language models may encounter factual knowledge during pre-training yet fail to reliably use that knowledge after fine-tuning. Despite g

QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2512.19134v2 Announce Type: replace Abstract: Dynamic Retrieval-Augmented Generation adaptively determines when to retrieve during generation to mitigate hallucinations in large language models

Query-Aware Learnable Graph Pooling Tokens as Prompt for Large Language Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2501.17549v2 Announce Type: replace Abstract: Graph-structured data plays a vital role in numerous domains, such as social networks, citation networks, commonsense reasoning graphs and knowledge

Readers make targeted regressions to plausible errors in reanalysis of 'noisy-channel garden-path' sentences

ResearchDGX agent

arXiv:2605.18563v1 Announce Type: new Abstract: A key question in psycholinguistics is how inferences about the meaning of linguistic input unfold incrementally a comprehender's mind. In this work, we

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts

Model ReleasesDGX agent

arXiv:2510.07239v2 Announce Type: replace Abstract: Automated red-teaming has emerged as a scalable approach for auditing Large Language Models (LLMs) prior to deployment, yet existing approaches lack

Residual Semantic Decomposition of Word Embeddings

Model ReleasesDGX agent

arXiv:2605.17482v1 Announce Type: new Abstract: We introduce Residual Semantic Decomposition (RSD), a neural additive decomposition of word embeddings that balances embedding reconstruction with relat

Responsible Federated LLMs via Safety Filtering and Constitutional AI

SafetyDGX agent

arXiv:2502.16691v2 Announce Type: replace Abstract: Recent research has increasingly focused on training large language models (LLMs) using federated learning, known as FedLLM. However, responsible AI

Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models

ResearchDGX agent

arXiv:2508.06974v2 Announce Type: replace Abstract: 1-bit LLM quantization offers significant advantages in reducing storage and computational costs. However, existing methods typically train 1-bit LL

Rethinking Table Pruning in TableQA: From Sequential Revisions to Gold Trajectory-Supervised Parallel Search

ResearchDGX agent

arXiv:2601.03851v2 Announce Type: replace Abstract: Table Question Answering (TableQA) benefits significantly from table pruning, which extracts compact sub-tables by eliminating redundant cells to st

Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free

Model ReleasesDGX agent

arXiv:2605.16767v1 Announce Type: new Abstract: Multi-label legal annotation requires assigning multiple labels from large, evolving taxonomies to long, fact-intensive documents, often under limited s

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers

ResearchDGX agent

arXiv:2605.16941v1 Announce Type: new Abstract: Diffusion Large Language Models (DLLMs) promise fast parallel generation, yet open-source DLLMs still face a severe quality-speed trade-off: acceleratin

RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis

Model ReleasesDGX agent

arXiv:2605.16843v1 Announce Type: new Abstract: India's Right to Information Act, 2005 gives every citizen the right to demand information from public authorities, yet in practice most people cannot m

SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening

Model ReleasesDGX agent

arXiv:2605.17610v1 Announce Type: cross Abstract: The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deplo

Scale Determines Whether Language Models Organize Representation Geometry for Prediction

SafetyDGX agent

arXiv:2605.17084v1 Announce Type: cross Abstract: In language models, what a representation encodes is determined by the geometry of its representation space: distances, not activations, carry meaning

Scaling Accessible Mathematics on arXiv: HTML Conversion and MathML 4

ResearchDGX agent

arXiv:2605.16562v1 Announce Type: new Abstract: We report on the ongoing development of arXiv's HTML Papers offering, available on every new TeX/LaTeX submission since its initial release in 2023. The

Scaling Laws for Code: A More Data-Hungry Regime

Model ReleasesDGX agent

arXiv:2510.08702v2 Announce Type: replace Abstract: Code Large Language Models (LLMs) are revolutionizing software engineering. However, scaling laws that guide the efficient training are predominantl

SEDD: Scalable and Efficient Dataset Deduplication with GPUs

Model ReleasesDGX agent

arXiv:2501.01046v4 Announce Type: replace Abstract: Dataset deduplication is widely recognized as a crucial preprocessing step that enhances data quality and improves the performance of large language

Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback

Model ReleasesDGX agent

arXiv:2605.17448v1 Announce Type: cross Abstract: Computer-aided design (CAD) is the backbone of modern industrial design, yet learned CAD generators still fall short of real engineering pipelines: th

Semantic Reranking at Inference Time for Hard Examples in Rhetorical Role Labeling

ApplicationsDGX agent

arXiv:2605.18007v1 Announce Type: new Abstract: Rhetorical Role Labeling (RRL) assigns a functional role to each sentence in a document and is widely used in legal, medical, and scientific domains. Wh

SIREM: Speech-Informed MRI Reconstruction with Learned Sampling

Model ReleasesDGX agent

arXiv:2605.18221v1 Announce Type: cross Abstract: Real-time magnetic resonance imaging (rtMRI) of speech production enables non-invasive visualization of dynamic vocal-tract motion and is valuable for

Sometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation

Model ReleasesDGX agent

arXiv:2605.17710v1 Announce Type: new Abstract: Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind h

Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs

ResearchDGX agent

arXiv:2505.19155v2 Announce Type: replace-cross Abstract: Due to the auto-regressive nature of current video large language models (Video-LLMs), the inference latency increases as the input sequence l

Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias

SafetyDGX agent

arXiv:2509.22061v2 Announce Type: replace-cross Abstract: Speech Continuation (SC) is the task of generating a coherent extension of a spoken prompt while preserving both semantic context and speaker

Spherical Steering: Geometry-Aware Activation Rotation for Language Models

ResearchDGX agent

arXiv:2602.08169v2 Announce Type: replace-cross Abstract: Inference-time steering offers a promising way to control language models (LMs) without retraining. However, standard approaches typically rel

Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models

SafetyDGX agent

arXiv:2605.17672v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance by generating long chains of thought (CoT), but often overthink, continuing to reason after a s

T-FIX: Text-Based Explanations with Features Interpretable to eXperts

SafetyDGX agent

arXiv:2511.04070v3 Announce Type: replace Abstract: As LLMs are deployed in knowledge-intensive settings (e.g., surgery, astronomy, therapy), users are often domain experts who expect not just answers

Taming 'Zombie'' Agents: A Markov State-Aware Framework for Resilient Multi-Agent Evolution

SafetyDGX agent

arXiv:2605.17348v1 Announce Type: new Abstract: Recent advancements in LLM-based multi-agent systems have demonstrated remarkable collaborative capabilities across complex tasks. To improve overall ef

Temporal Decay of Co-Citation Predictability: A 20-Year Statute Retrieval Benchmark from 396M Ukrainian Court Citations

Model ReleasesDGX agent

arXiv:2605.17639v1 Announce Type: new Abstract: Co-citation structure is widely assumed to provide stable retrieval signal in legal information systems. We test this assumption longitudinally by const

The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought

Model ReleasesDGX agent

arXiv:2605.18079v1 Announce Type: cross Abstract: Existing expressivity results for transformers typically rely on hardmax attention, high precision, and other architectural modifications that disconn

The Frequency Confound in Language-Model Surprisal and Metaphor Novelty

ResearchDGX agent

arXiv:2605.06506v2 Announce Type: replace Abstract: Language-model (LM) surprisal is widely used as a proxy for contextual predictability and has been reported to correlate with metaphor novelty judgm

The Unlearnability Phenomenon in RLVR for Language Models

ResearchDGX agent

arXiv:2605.16787v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Reward (RLVR) has proven effective in improving Large Language Model's (LLM) reasoning ability. However, the le

To MRL or not to MRL: Text Embeddings are Robust to Truncation Without Matryoshka Embeddings, Except In Heavy Truncation Scenarios

ResearchDGX agent

arXiv:2605.16608v1 Announce Type: cross Abstract: Matryoshka Representation Learning (MRL) is a widely adopted approach for training text encoders so they provide useful text representations at variou

ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints

Model ReleasesDGX agent

arXiv:2602.21265v2 Announce Type: replace Abstract: We introduce ToolMATH, a math-grounded diagnostic benchmark for evaluating long-horizon tool use under controllable tool-catalog conditions. ToolMAT

Traces of Social Competence in Large Language Models

ApplicationsDGX agent

arXiv:2603.04161v2 Announce Type: replace Abstract: The False Belief Test (FBT) has been the main method for assessing Theory of Mind (ToM) and related socio-cognitive competencies. For Large Language

Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

Model ReleasesDGX agent

arXiv:2605.17453v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly rely on external tools to make consequential decisions, yet most existing agent-security benchmarks and defenses im

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

Model ReleasesDGX agent

arXiv:2601.12696v2 Announce Type: replace Abstract: Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerab

Universal Adversarial Triggers

ResearchDGX agent

arXiv:2605.17936v1 Announce Type: new Abstract: Recent works have illustrated that modern NLP models trained for diverse tasks ranging from sentiment analysis to language generation succumb to univers

Vector RAG vs LLM-Compiled Wiki: A Preregistered Comparison on a Small Multi-Domain Research

ResearchDGX agent

arXiv:2605.18490v1 Announce Type: new Abstract: We preregistered a comparison of two ways to help an LLM answer questions over a small research corpus: a single-round Vector RAG system and an LLM-comp

VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.17467v1 Announce Type: new Abstract: Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliabil

Vidya: An AI-Driven Modular Pipeline for Archival Automation and Semantic Metadata Enrichment

ResearchDGX agent

arXiv:2605.16338v1 Announce Type: cross Abstract: The large-scale digitization of historical archives has created a paradox: 'dark data'-digital objects lacking metadata for retrieval. Manual archival

We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong

Model ReleasesDGX agent

arXiv:2509.22510v3 Announce Type: replace Abstract: Alignment of Large Language Models (LLMs) is the ability to satisfy desired objectives during generation, which is critical for trustworthy deployme

WEBSERV: A Full-Stack and RL-Ready Web Environment for Training Web Agents at Scale

Model ReleasesDGX agent

arXiv:2510.16252v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) for web agents demands environments that are both effective for evaluation and efficient enough for large-scale on

When AI Tells You What You Want to Hear: Sycophantic Behavior of Large Language Models in Dementia Care Settings

Model ReleasesDGX agent

arXiv:2605.16288v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in clinical and care settings. This exploratory study investigates whether LLMs exhibit sycophantic

When TableQA Meets Noise: A Dual Denoising Framework for Complex Questions and Large-scale Tables

TutorialsDGX agent

arXiv:2509.17680v2 Announce Type: replace Abstract: Table question answering (TableQA) is a fundamental task in natural language processing (NLP). The strong reasoning capabilities of large language m

White-Box Sensitivity Auditing with Steering Vectors

SafetyDGX agent

arXiv:2601.16398v2 Announce Type: replace-cross Abstract: Algorithmic audits are essential tools for examining systems for properties required by regulators or desired by operators. Current audits of

You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations

ResearchDGX agent

arXiv:2511.06516v3 Announce Type: replace Abstract: Many LLM applications require only narrow capabilities, yet standard post-training quantization (PTQ) methods allocate precision without considering

18 May 2026

Adesua: Development and Feasibility Study of an AI WhatsApp Bot for Science Learning in West Africa

ApplicationsDGX agent

arXiv:2605.15376v1 Announce Type: new Abstract: Sub-Saharan Africa faces persistently high student-teacher ratios and shortages of qualified teachers, limiting students' access to personalized learnin

AirNav: A Large-Scale UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions

Model ReleasesDGX agent

arXiv:2601.03707v2 Announce Type: replace Abstract: Existing UAV vision-and-language navigation (VLN) benchmarks rarely provide realistic aerial scenes, natural process-level instructions, and suffici

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination

ResearchDGX agent

arXiv:2605.15864v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) often produce self-reflective statements like 'let me check the figure again' during reasoning. Do such statements trigg

Artificial Aphasias in Lesioned Language Models

ResearchDGX agent

arXiv:2605.16222v1 Announce Type: new Abstract: Aphasias, selective language impairments which can arise from brain damage, reveal the functional organization of human language by providing causal lin

Automatic Construction of a Legal Citation Graph from 100 Million Ukrainian Court Decisions: Large-Scale Extraction, Topological Analysis, and Ontology-Driven Clustering

Model ReleasesDGX agent

arXiv:2605.15362v1 Announce Type: new Abstract: Half a billion citation edges extracted from 100.7 million Ukrainian court decisions reveal that judicial citation structure encodes legal domain bounda

Beyond Forgetting: Machine Unlearning Elicits Controllable Side Behaviors and Capabilities

ResearchDGX agent

arXiv:2601.21702v3 Announce Type: replace-cross Abstract: We consider Representation Misdirection (RM), a class of large language model (LLM) unlearning methods that achieve forgetting by redirecting

BootstrapAgent: Distilling Repository Setup into Reusable Agent Knowledge

AgentsDGX agent

arXiv:2605.15815v1 Announce Type: cross Abstract: Code agents increasingly help developers work with unfamiliar repositories, but every such task depends on a costly prerequisite: bootstrapping the re

Calibrating LLMs with Semantic-level Reward

ApplicationsDGX agent

arXiv:2605.15588v1 Announce Type: new Abstract: As large language models (LLMs) are deployed in consequential settings such as medical question answering and legal reasoning, the ability to estimate w

Can Large Language Models Imitate Human Speech for Clinical Assessment? LLM-Driven Data Augmentation for Cognitive Score Prediction

Model ReleasesDGX agent

arXiv:2605.16077v1 Announce Type: new Abstract: Accurate assessment of cognitive decline from spontaneous speech remains challenging due to limited dataset size and class imbalance. In this work, we p

Capability Conditioned Scaffolding for Professional Human LLM Collaboration

ResearchDGX agent

arXiv:2605.15404v1 Announce Type: new Abstract: Large language model personalization typically adapts outputs to user preferences and style but does not account for differences in user evaluation capa

Contexting as Recommendation: Evolutionary Collaborative Filtering for Context Engineering

TutorialsDGX agent

arXiv:2605.15721v1 Announce Type: new Abstract: Large Language Models (LLMs) are highly sensitive to their input contexts, motivating the development of automated context engineering. However, existin

Conversations in Space: Structuring Non-Linear LLM Interactions on a Canvas

ResearchDGX agent

arXiv:2605.15848v1 Announce Type: cross Abstract: Conversational interfaces powered by large language models (LLMs) are widely used for ideation and analysis, yet their linear structure limits explora

CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency

Model ReleasesDGX agent

arXiv:2512.00417v5 Announce Type: replace Abstract: This paper introduces CryptoBench, the first expert-curated, dynamic benchmark designed to rigorously evaluate the real-world capabilities of Large

Defining Cultural Capabilities for AI Evaluation: A Taxonomy Grounded in Intercultural Communication Theory

ApplicationsDGX agent

arXiv:2605.15990v1 Announce Type: new Abstract: Tremendous efforts have been put into evaluating the inclusivity and effectiveness of AI systems across cultures. However, the cultural capabilities con

← Previous
1…7172737475…129
Next →