AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
12 May 2026

Machine Learning Research Has Outpaced Its Communication Norms and NeurIPS Should Act

ResearchDGX agent

arXiv:2605.08889v1 Announce Type: cross Abstract: Machine learning research has grown exponentially while its communication norms have not. We argue NeurIPS should adopt explicit, measurable writing s

Majority Bit-Aware Watermarking For Large Language Models

ResearchDGX agent

arXiv:2508.03829v2 Announce Type: replace Abstract: The growing deployment of Large Language Models (LLMs) has raised concerns about their misuse in generating harmful or deceptive content. To address

MapFormer: Self-Supervised Learning of Cognitive Maps with Input-Dependent Positional Embeddings

SafetyDGX agent

arXiv:2511.19279v4 Announce Type: replace-cross Abstract: A cognitive map is an internal model which encodes the abstract relationships among entities in the world, giving humans and animals the flexi


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MARS-SQL: A multi-agent reinforcement learning framework for Text-to-SQL

SafetyDGX agent

arXiv:2511.01008v2 Announce Type: replace Abstract: Large Language Models (LLMs) often struggle with the precise logic and schema alignment required for complex Text-to-SQL tasks. While current method

Max-pooling Network Revisited: Analyzing the Role of Semantic Probability in Multiple Instance Learning for Hallucination Detection

ResearchDGX agent

arXiv:2605.08863v1 Announce Type: new Abstract: Hallucination detection has become increasingly important for improving the reliability of large language models (LLMs). Recently, hybrid approaches suc

Mela: Test-Time Memory Consolidation based on Transformation Hypothesis

ResearchDGX agent

arXiv:2605.10537v1 Announce Type: new Abstract: Memory consolidation, the process by which transient experiences are transformed into stable, structured representations, is a foundational organizing p

Mem-W: Latent Memory-Native GUI Agents

SafetyDGX agent

arXiv:2605.09317v1 Announce Type: new Abstract: GUI agents are beginning to operate the web, mobile, and desktop as interactive worlds, where successful control depends on carrying forward visual, pro

MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents

Model ReleasesDGX agent

arXiv:2605.09530v1 Announce Type: cross Abstract: As LLM-powered agents are increasingly deployed in edge-cloud environments, personalized memory has become a key enabler of long-term adaptation and u

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology

Model ReleasesDGX agent

arXiv:2605.09152v1 Announce Type: new Abstract: Deciphering animal intent is a fundamental challenge in computational ethology, largely because of semantic aliasing, the phenomenon where identical ext

Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inference

Local AiDGX agent

arXiv:2605.09990v1 Announce Type: new Abstract: Data-intensive applications, ranging from large-scale retrieval systems to advanced data pipelines, are increasingly bottlenecked by the processing of h

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

Model ReleasesDGX agent

arXiv:2510.09592v2 Announce Type: replace Abstract: Real-time Spoken Language Models (SLMs) struggle to leverage Chain-of-Thought (CoT) reasoning due to the prohibitive latency of generating the entir

Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer

ResearchDGX agent

arXiv:2602.12286v2 Announce Type: replace-cross Abstract: A central challenge in developing Multimodal Large Language Models (MLLMs) is effectively integrating heterogeneous inputs into a cohesive rea

Model-Aware Tokenizer Transfer

HardwareDGX agent

arXiv:2510.21954v2 Announce Type: replace Abstract: Large Language Models (LLMs) are trained to support an increasing number of languages, yet their predefined tokenizers remain a bottleneck for adapt

MolSight: Molecular Property Prediction with Images

Model ReleasesDGX agent

arXiv:2605.10157v1 Announce Type: cross Abstract: Every molecule ever synthesised can be drawn as a 2D skeletal diagram, yet in modern property prediction this universally available representation has

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

Model ReleasesDGX agent

arXiv:2605.10616v1 Announce Type: cross Abstract: Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generaliza

Multi-domain Multi-modal Document Classification Benchmark with a Multi-level Taxonomy

Model ReleasesDGX agent

arXiv:2605.10550v1 Announce Type: new Abstract: Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -

MUR: Momentum Uncertainty guided Reasoning

TutorialsDGX agent

arXiv:2507.14958v2 Announce Type: replace Abstract: Current models have achieved impressive performance on reasoning-intensive tasks, yet optimizing their reasoning efficiency remains an open challeng

NARRA-Gym for Evaluating Interactive Narrative Agents

Model ReleasesDGX agent

arXiv:2605.08503v1 Announce Type: new Abstract: Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmark

Nectar: Neural Estimation of Cached-Token Attention via Regression

ApplicationsDGX agent

arXiv:2605.09778v1 Announce Type: cross Abstract: Evaluating softmax attention over a fixed long context requires reading every cached key-value pair for each new query token. For a given context (a b

Neural at ArchEHR-QA 2026: One Method Fits All: Unified Prompt Optimization for Clinical QA over EHRs

SafetyDGX agent

arXiv:2605.10877v1 Announce Type: new Abstract: Automated question answering (QA) over electronic health records (EHRs) demands precise evidence retrieval, faithful answer generation, and explicit gro

Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes

HardwareDGX agent

arXiv:2605.08913v1 Announce Type: cross Abstract: Autoregressive inference is typically assumed to scale predictably with decoding length, and key-value (KV) caching is widely regarded as a universall

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness

Model ReleasesDGX agent

arXiv:2605.10379v1 Announce Type: new Abstract: Large language models (LLMs) have become capable mathematical problem-solvers, often producing correct proofs for challenging problems. However, correct

Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning

HardwareDGX agent

arXiv:2605.09490v1 Announce Type: new Abstract: Reasoning LLMs produce thousands of chain-of-thought tokens whose KV cache must reside in scarce GPU HBM. The dominant response -- permanently evicting

NyayaAI: An AI-Powered Legal Assistant Using Multi-Agent Architecture and Retrieval-Augmented Generation

AgentsDGX agent

arXiv:2605.10155v1 Announce Type: new Abstract: Legal information in India remains largely inaccessible due to the complexity of legal language and the sheer volume of legal documentation involved in

Overconfident and Blind to Details: Fixing Prompt Insensitivity with Abductive Preference Learning

Model ReleasesDGX agent

arXiv:2510.09887v2 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pretraining priors. For example, models will confident

PAAC: Privacy-Aware Agentic Device-Cloud Collaboration

Local AiDGX agent

arXiv:2605.08646v1 Announce Type: cross Abstract: Large language model (LLM) agents face a structural tension: cloud agents provide strong reasoning but expose user data, while on-device agents preser

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs

SafetyDGX agent

arXiv:2605.09422v1 Announce Type: new Abstract: Although Large Multimodal Models (LMMs) have achieved strong performance on general video understanding, their susceptibility to textual prior shortcuts

PHAGE: Patent Heterogeneous Attention-Guided Graph Encoder for Representation Learning

SafetyDGX agent

arXiv:2605.10073v1 Announce Type: new Abstract: Patent claims form a directed dependency structure in which dependent claims inherit and refine the scope of earlier claims; however, existing patent en

PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoning

Model ReleasesDGX agent

arXiv:2605.10032v1 Announce Type: new Abstract: Cell-type-specific marker genes are fundamental to plant biology, yet existing resources primarily rely on curated databases or high-throughput studies

PumpSense: Real-Time Detection and Target Extraction of Crypto Pump-and-Dumps on Telegram

Model ReleasesDGX agent

arXiv:2605.09431v1 Announce Type: new Abstract: Cryptocurrency pump-and-dump schemes coordinated via Telegram threaten market integrity. However, existing research addressing this specific threat has

QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model

Model ReleasesDGX agent

arXiv:2504.12334v2 Announce Type: replace Abstract: Large language models (LLMs) face significant challenges in specialized biomedical tasks due to the inherent complexity of medical reasoning and the

Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants

Model ReleasesDGX agent

arXiv:2605.09808v1 Announce Type: new Abstract: User simulators are increasingly leveraged to build interactive AI assistants, yet how to measure the quality of these simulators remains an open questi

Queryable LoRA: Instruction-Regularized Routing Over Shared Low-Rank Update Atoms

Model ReleasesDGX agent

arXiv:2605.08423v1 Announce Type: cross Abstract: We present a data-adaptive method for parameter-efficient fine-tuning of large neural networks. Standard low-rank adaptation methods improve efficienc

Reasoning emerges from constrained inference manifolds in large language models

Model ReleasesDGX agent

arXiv:2605.08142v1 Announce Type: cross Abstract: Reasoning in large language models is predominantly evaluated through labeled benchmarks, conflating task performance with the quality of internal inf

Reasoning Trajectories for Socratic Debugging of Student Code: From Misconceptions to Contradictions and Updated Beliefs

TutorialsDGX agent

arXiv:2511.00371v2 Announce Type: replace Abstract: In Socratic debugging, instructors guide students towards identifying and fixing a bug on their own, instead of providing the bug fix directly. Most

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR

ResearchDGX agent

arXiv:2605.10781v1 Announce Type: cross Abstract: Self-distillation has emerged as a powerful framework for post-training LLMs, where a teacher conditioned on extra information guides a student withou

Reinforcing Multimodal Reasoning Against Visual Degradation

SafetyDGX agent

arXiv:2605.09262v1 Announce Type: cross Abstract: Reinforcement Learning has significantly advanced the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet the resulting policies r

Relative Kinetic Utility for Reasoning-Aware Structural Pruning in Large Language Models

Model ReleasesDGX agent

arXiv:2605.09008v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) prompting symbolized a huge improvement of reasoning capabilities of Large Language Models (LLMs). However, scaling up test-tim

Relative Score Policy Optimization for Diffusion Language Models

SafetyDGX agent

arXiv:2605.10218v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) offer a promising route to parallel and efficient text generation, but improving their reasoning ability require

Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs

Model ReleasesDGX agent

arXiv:2605.09239v1 Announce Type: new Abstract: Large language models fail at counting repeated tokens despite strong performance on broader reasoning benchmarks. These failures are commonly attribute

Responsible Benchmarking of Fairness for Automatic Speech Recognition

SafetyDGX agent

arXiv:2605.10615v1 Announce Type: new Abstract: Many studies have shown automatic speech processing (ASR) systems have unequal performance across speakergroups (SG's). However, the manner in which suc

ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal Smoothing

ResearchDGX agent

arXiv:2605.08840v1 Announce Type: new Abstract: Large language models (LLMs) face growing challenges in efficient generative inference due to the increasing memory demands of Key-Value (KV) caches, es

Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models

ResearchDGX agent

arXiv:2602.01698v3 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-tra

Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning

ResearchDGX agent

arXiv:2512.11470v2 Announce Type: replace-cross Abstract: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) dominate the post-training landscape for mathematical reasoning, yet differ funda

Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning

SafetyDGX agent

arXiv:2605.06241v2 Announce Type: replace Abstract: Reinforcement learning has become the standard for improving reasoning in large language models, yet evidence increasingly suggests that RL does not

Revisiting the syntax of imperatives in Yemeni Arabic: An Agree across phases approach

ResearchDGX agent

arXiv:2605.08447v1 Announce Type: new Abstract: This article revisits the syntax of imperatives in Yemeni Arabic proposing an Agree acros phases (AAP) approach. I argue that the AAP approach successfu

Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios

SafetyDGX agent

arXiv:2512.00920v4 Announce Type: replace Abstract: Reliable reward models (RMs) are critical for ensuring the safe alignment of large language models (LLMs). However, current RM evaluation methods fo

Route Before Retrieve: Activating Latent Routing Abilities of LLMs for RAG vs. Long-Context Selection

Model ReleasesDGX agent

arXiv:2605.10235v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have expanded the context window to beyond 128K tokens, enabling long-document understanding and multi-s

RUBEN: Rule-Based Explanations for Retrieval-Augmented LLM Systems

SafetyDGX agent

arXiv:2605.10862v1 Announce Type: new Abstract: This paper demonstrates RUBEN, an interactive tool for discovering minimal rules to explain the outputs of retrieval-augmented large language models (LL

RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards

SafetyDGX agent

arXiv:2605.10899v1 Announce Type: new Abstract: Training deep research agents, namely systems that plan, search, evaluate evidence, and synthesize long-form reports, pushes reinforcement learning beyo

SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators

SafetyDGX agent

arXiv:2605.08334v1 Announce Type: new Abstract: We present SalesSim, a framework and testbed for evaluating the ability of Multimodal Large Language Models (MLLMs) to simulate realistic, persona-drive

Scaling Mobile Agent Systems: From Capability Density to Collective Intelligence

Local AiDGX agent

arXiv:2605.08124v1 Announce Type: cross Abstract: Mobile agent systems are emerging as a key paradigm for enabling intelligent applications on edge devices and in AIoT ecosystems. However, their scala

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models

ResearchDGX agent

arXiv:2605.09630v1 Announce Type: new Abstract: Tokenizer-free language models eliminate the tokenizer step of the language modeling pipeline by operating directly on bytes; patch-based variants furth

SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness

Model ReleasesDGX agent

arXiv:2603.14889v2 Announce Type: replace-cross Abstract: The rapid evolution of end-to-end spoken dialogue systems demands transcending mere textual semantics to incorporate paralinguistic nuances an

SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization

AgentsDGX agent

arXiv:2605.08382v1 Announce Type: cross Abstract: LLM coding agents now generate code at an unprecedented scale, yet LLM-generated code introduces cybersecurity vulnerabilities into codebases without

ShifaMind: A Multiplicative Concept Bottleneck for Interpretable ICD-10 Coding

ResearchDGX agent

arXiv:2605.08482v1 Announce Type: cross Abstract: Automated ICD-10 coding from clinical discharge summaries requires models that are both accurate on long-tailed multi-label classification tasks and i

SkillMAS: Skill Co-Evolution with LLM-based Multi-Agent System

AgentsDGX agent

arXiv:2605.09341v1 Announce Type: cross Abstract: Large language model (LLM) agent systems are increasingly expected to improve after deployment, but existing work often decouples two adaptation targe

SkillRAE: Agent Skill-Based Context Compilation for Retrieval-Augmented Execution

AgentsDGX agent

arXiv:2605.10114v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents (e.g., OpenClaw) increasingly rely on reusable skill libraries to solve artifact-rich tasks such as document-cen

SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decoding

ResearchDGX agent

arXiv:2605.10453v1 Announce Type: cross Abstract: Speculative decoding speeds up autoregressive generation in Large Language Models (LLMs) through a two-step procedure, where a lightweight draft model

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

Model ReleasesDGX agent

arXiv:2605.09063v1 Announce Type: new Abstract: Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challengi

← Previous
1…7980818283…129
Next →