AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlog
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Safety

Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics

DGX agent

arXiv:2605.18549v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not alwa

safetyarxiv-cs-cl
19 May 2026
Safety
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages

DGX agent

arXiv:2605.17152v1 Announce Type: new Abstract: Multimodal LLMs are evolving from vision-language to tri-modality that see, hear, and read, yet pipelines and benchmarks remain English-centric and comp

safetyarxiv-cs-cl
19 May 2026
Model Releases

Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Models

DGX agent

arXiv:2605.16409v1 Announce Type: cross Abstract: Optical character recognition (OCR) and multilingual text understanding remain major failure modes of multimodal large language models (MLLMs), partic

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

NewsLens: A Multi-Agent Framework for Adversarial News Bias Navigation

DGX agent

arXiv:2605.17364v1 Announce Type: new Abstract: Media bias detection has predominantly been framed as a classification task: assign a political label to an article or outlet. We argue this framing is

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

PaliBench: A Multi-Reference Blueprint for Classical Language Translation Benchmarks

DGX agent

arXiv:2605.16881v1 Announce Type: new Abstract: Digital humanities projects increasingly rely on machine translation and large language models to widen access to classical, religious, and otherwise un

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

PEGRL: Improving Machine Translation by Post-Editing Guided Reinforcement Learning

DGX agent

arXiv:2602.03352v2 Announce Type: replace Abstract: Reinforcement learning (RL) has shown strong promise for LLM-based machine translation, with recent methods such as GRPO demonstrating notable gains

model-releasesarxiv-cs-cl
19 May 2026
Local Ai

PPAI: Enabling Personalized LLM Agent Interoperability for Collaborative Edge Intelligence

DGX agent

arXiv:2605.18067v1 Announce Type: new Abstract: Deploying large language model (LLM) on edge device enables personalized LLM agents for various users. The growing availability of diverse personalized

local-aiarxiv-cs-cl
19 May 2026
Safety

PQR: A Framework to Generate Diverse and Realistic User Queries that Elicit QA Agent Failures

DGX agent

arXiv:2605.16551v1 Announce Type: new Abstract: Evaluating LLM-based agents remains challenging because identifying meaningful failure cases often requires substantial human effort to design realistic

safetyarxiv-cs-cl
19 May 2026
Safety

Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs

DGX agent

arXiv:2605.18352v1 Announce Type: new Abstract: Presupposition projection in conditionals is central to theories of meaning and pragmatics, yet it remains largely unevaluated in large language models.

safetyarxiv-cs-cl
19 May 2026
Agents

Proof-Carrying Certificates for LLM Pipelines: A Trust-Boundary Architecture

DGX agent

arXiv:2605.16407v1 Announce Type: cross Abstract: We present a framework for verifying the deterministic structured computations surrounding a large language model rather than the model itself, extend

agentsarxiv-cs-cl
19 May 2026
Model Releases

Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction

DGX agent

arXiv:2605.18053v1 Announce Type: cross Abstract: We study KV cache eviction under a shared globally capped decode-time harness. Seven policies (LRU, H2O, SnapKV, StreamingLLM, Ada-KV, QUEST, Random)

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Provable Knowledge Acquisition and Extraction in One-Layer Transformers

DGX agent

arXiv:2508.00901v4 Announce Type: replace-cross Abstract: Large language models may encounter factual knowledge during pre-training yet fail to reliably use that knowledge after fine-tuning. Despite g

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation

DGX agent

arXiv:2512.19134v2 Announce Type: replace Abstract: Dynamic Retrieval-Augmented Generation adaptively determines when to retrieve during generation to mitigate hallucinations in large language models

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Query-Aware Learnable Graph Pooling Tokens as Prompt for Large Language Models

DGX agent

arXiv:2501.17549v2 Announce Type: replace Abstract: Graph-structured data plays a vital role in numerous domains, such as social networks, citation networks, commonsense reasoning graphs and knowledge

model-releasesarxiv-cs-cl
19 May 2026
Research

Readers make targeted regressions to plausible errors in reanalysis of 'noisy-channel garden-path' sentences

DGX agent

arXiv:2605.18563v1 Announce Type: new Abstract: A key question in psycholinguistics is how inferences about the meaning of linguistic input unfold incrementally a comprehender's mind. In this work, we

researcharxiv-cs-cl
19 May 2026
Model Releases

Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA Experts

DGX agent

arXiv:2510.07239v2 Announce Type: replace Abstract: Automated red-teaming has emerged as a scalable approach for auditing Large Language Models (LLMs) prior to deployment, yet existing approaches lack

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Residual Semantic Decomposition of Word Embeddings

DGX agent

arXiv:2605.17482v1 Announce Type: new Abstract: We introduce Residual Semantic Decomposition (RSD), a neural additive decomposition of word embeddings that balances embedding reconstruction with relat

model-releasesarxiv-cs-cl
19 May 2026
Safety

Responsible Federated LLMs via Safety Filtering and Constitutional AI

DGX agent

arXiv:2502.16691v2 Announce Type: replace Abstract: Recent research has increasingly focused on training large language models (LLMs) using federated learning, known as FedLLM. However, responsible AI

safetyarxiv-cs-cl
19 May 2026
Research

Rethinking 1-bit Optimization Leveraging Pre-trained Large Language Models

DGX agent

arXiv:2508.06974v2 Announce Type: replace Abstract: 1-bit LLM quantization offers significant advantages in reducing storage and computational costs. However, existing methods typically train 1-bit LL

researcharxiv-cs-cl
19 May 2026
Research

Rethinking Table Pruning in TableQA: From Sequential Revisions to Gold Trajectory-Supervised Parallel Search

DGX agent

arXiv:2601.03851v2 Announce Type: replace Abstract: Table Question Answering (TableQA) benefits significantly from table pruning, which extracts compact sub-tables by eliminating redundant cells to st

researcharxiv-cs-cl
19 May 2026
Model Releases

Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free

DGX agent

arXiv:2605.16767v1 Announce Type: new Abstract: Multi-label legal annotation requires assigning multiple labels from large, evolving taxonomies to long, fact-intensive documents, often under limited s

model-releasesarxiv-cs-cl
19 May 2026
Research

Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachers

DGX agent

arXiv:2605.16941v1 Announce Type: new Abstract: Diffusion Large Language Models (DLLMs) promise fast parallel generation, yet open-source DLLMs still face a severe quality-speed trade-off: acceleratin

researcharxiv-cs-cl
19 May 2026
Model Releases

RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis

DGX agent

arXiv:2605.16843v1 Announce Type: new Abstract: India's Right to Information Act, 2005 gives every citizen the right to demand information from public authorities, yet in practice most people cannot m

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening

DGX agent

arXiv:2605.17610v1 Announce Type: cross Abstract: The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deplo

model-releasesarxiv-cs-cl
19 May 2026
Safety

Scale Determines Whether Language Models Organize Representation Geometry for Prediction

DGX agent

arXiv:2605.17084v1 Announce Type: cross Abstract: In language models, what a representation encodes is determined by the geometry of its representation space: distances, not activations, carry meaning

safetyarxiv-cs-cl
19 May 2026
Research

Scaling Accessible Mathematics on arXiv: HTML Conversion and MathML 4

DGX agent

arXiv:2605.16562v1 Announce Type: new Abstract: We report on the ongoing development of arXiv's HTML Papers offering, available on every new TeX/LaTeX submission since its initial release in 2023. The

researcharxiv-cs-cl
19 May 2026
Model Releases

Scaling Laws for Code: A More Data-Hungry Regime

DGX agent

arXiv:2510.08702v2 Announce Type: replace Abstract: Code Large Language Models (LLMs) are revolutionizing software engineering. However, scaling laws that guide the efficient training are predominantl

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

SEDD: Scalable and Efficient Dataset Deduplication with GPUs

DGX agent

arXiv:2501.01046v4 Announce Type: replace Abstract: Dataset deduplication is widely recognized as a crucial preprocessing step that enhances data quality and improves the performance of large language

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback

DGX agent

arXiv:2605.17448v1 Announce Type: cross Abstract: Computer-aided design (CAD) is the backbone of modern industrial design, yet learned CAD generators still fall short of real engineering pipelines: th

model-releasesarxiv-cs-cl
19 May 2026
Applications

Semantic Reranking at Inference Time for Hard Examples in Rhetorical Role Labeling

DGX agent

arXiv:2605.18007v1 Announce Type: new Abstract: Rhetorical Role Labeling (RRL) assigns a functional role to each sentence in a document and is widely used in legal, medical, and scientific domains. Wh

applicationsarxiv-cs-cl
19 May 2026
Model Releases

SIREM: Speech-Informed MRI Reconstruction with Learned Sampling

DGX agent

arXiv:2605.18221v1 Announce Type: cross Abstract: Real-time magnetic resonance imaging (rtMRI) of speech production enables non-invasive visualization of dynamic vocal-tract motion and is valuable for

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

Sometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation

DGX agent

arXiv:2605.17710v1 Announce Type: new Abstract: Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind h

model-releasesarxiv-cs-cl
19 May 2026
Research

Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs

DGX agent

arXiv:2505.19155v2 Announce Type: replace-cross Abstract: Due to the auto-regressive nature of current video large language models (Video-LLMs), the inference latency increases as the input sequence l

researcharxiv-cs-cl
19 May 2026
Safety

Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias

DGX agent

arXiv:2509.22061v2 Announce Type: replace-cross Abstract: Speech Continuation (SC) is the task of generating a coherent extension of a spoken prompt while preserving both semantic context and speaker

safetyarxiv-cs-cl
19 May 2026
Research

Spherical Steering: Geometry-Aware Activation Rotation for Language Models

DGX agent

arXiv:2602.08169v2 Announce Type: replace-cross Abstract: Inference-time steering offers a promising way to control language models (LMs) without retraining. However, standard approaches typically rel

researcharxiv-cs-cl
19 May 2026
Safety

Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models

DGX agent

arXiv:2605.17672v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance by generating long chains of thought (CoT), but often overthink, continuing to reason after a s

safetyarxiv-cs-cl
19 May 2026
Safety

T-FIX: Text-Based Explanations with Features Interpretable to eXperts

DGX agent

arXiv:2511.04070v3 Announce Type: replace Abstract: As LLMs are deployed in knowledge-intensive settings (e.g., surgery, astronomy, therapy), users are often domain experts who expect not just answers

safetyarxiv-cs-cl
19 May 2026
Safety

Taming 'Zombie'' Agents: A Markov State-Aware Framework for Resilient Multi-Agent Evolution

DGX agent

arXiv:2605.17348v1 Announce Type: new Abstract: Recent advancements in LLM-based multi-agent systems have demonstrated remarkable collaborative capabilities across complex tasks. To improve overall ef

safetyarxiv-cs-cl
19 May 2026
Model Releases

Temporal Decay of Co-Citation Predictability: A 20-Year Statute Retrieval Benchmark from 396M Ukrainian Court Citations

DGX agent

arXiv:2605.17639v1 Announce Type: new Abstract: Co-citation structure is widely assumed to provide stable retrieval signal in legal information systems. We test this assumption longitudinally by const

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

The Expressive Power of Low Precision Softmax Transformers with (Summarized) Chain-of-Thought

DGX agent

arXiv:2605.18079v1 Announce Type: cross Abstract: Existing expressivity results for transformers typically rely on hardmax attention, high precision, and other architectural modifications that disconn

model-releasesarxiv-cs-cl
19 May 2026
Research

The Frequency Confound in Language-Model Surprisal and Metaphor Novelty

DGX agent

arXiv:2605.06506v2 Announce Type: replace Abstract: Language-model (LM) surprisal is widely used as a proxy for contextual predictability and has been reported to correlate with metaphor novelty judgm

researcharxiv-cs-cl
19 May 2026
Research

The Unlearnability Phenomenon in RLVR for Language Models

DGX agent

arXiv:2605.16787v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Reward (RLVR) has proven effective in improving Large Language Model's (LLM) reasoning ability. However, the le

researcharxiv-cs-cl
19 May 2026
Research

To MRL or not to MRL: Text Embeddings are Robust to Truncation Without Matryoshka Embeddings, Except In Heavy Truncation Scenarios

DGX agent

arXiv:2605.16608v1 Announce Type: cross Abstract: Matryoshka Representation Learning (MRL) is a widely adopted approach for training text encoders so they provide useful text representations at variou

researcharxiv-cs-cl
19 May 2026
Model Releases

ToolMATH: A Diagnostic Benchmark for Long-Horizon Tool Use under Systematic Tool-Catalog Constraints

DGX agent

arXiv:2602.21265v2 Announce Type: replace Abstract: We introduce ToolMATH, a math-grounded diagnostic benchmark for evaluating long-horizon tool use under controllable tool-catalog conditions. ToolMAT

model-releasesarxiv-cs-cl
19 May 2026
Applications

Traces of Social Competence in Large Language Models

DGX agent

arXiv:2603.04161v2 Announce Type: replace Abstract: The False Belief Test (FBT) has been the main method for assessing Theory of Mind (ToM) and related socio-cognitive competencies. For Large Language

applicationsarxiv-cs-cl
19 May 2026
Model Releases

Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedback

DGX agent

arXiv:2605.17453v1 Announce Type: cross Abstract: Tool-using LLM agents increasingly rely on external tools to make consequential decisions, yet most existing agent-security benchmarks and defenses im

model-releasesarxiv-cs-cl
19 May 2026
Model Releases

UbuntuGuard: A Culturally-Grounded Policy Benchmark for Equitable AI Safety in African Languages

DGX agent

arXiv:2601.12696v2 Announce Type: replace Abstract: Current guardian models are predominantly Western-centric and optimized for high-resource languages, leaving low-resource African languages vulnerab

model-releasesarxiv-cs-cl
19 May 2026
Research

Universal Adversarial Triggers

DGX agent

arXiv:2605.17936v1 Announce Type: new Abstract: Recent works have illustrated that modern NLP models trained for diverse tasks ranging from sentiment analysis to language generation succumb to univers

researcharxiv-cs-cl
19 May 2026
← Previous
1…9091929394…162
Next →