AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
24 Jul 2026

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

Model ReleasesDGX agent

arXiv:2607.20891v1 Announce Type: new Abstract: Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence synthesis, and report generation, y

Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought

Model ReleasesDGX agent

arXiv:2607.20427v1 Announce Type: cross Abstract: Mixture-of-Experts architectures have revolutionized scaling, yet the underlying logic of their routing remains a black box. In this paper, we uncover

Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.20494v1 Announce Type: new Abstract: Production LLM applications commonly stack a regex filter in front of model-side alignment; prior work found no measurable coverage gain from adding a l

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

Model ReleasesDGX agent

arXiv:2607.20759v1 Announce Type: cross Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with au

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

Model ReleasesDGX agent

arXiv:2607.20466v1 Announce Type: new Abstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equiv

Joint Utilization of Geospatial and census proxies for Autoencoder-Assisted Downscaling (JUGAAD) of socioeconomic indicators in India

ApplicationsDGX agent

arXiv:2607.20559v1 Announce Type: cross Abstract: Monitoring poverty and food security indicators is imperative for addressing socioeconomic challenges in developing nations. A limitation is mismatche

KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback

SafetyDGX agent

arXiv:2607.20556v1 Announce Type: new Abstract: In large-scale text analysis tasks, pre-trained language models are often used to embed text corpora for downstream analysis. However, such models may s

Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs'Hallucinations

ResearchDGX agent

arXiv:2607.20426v1 Announce Type: cross Abstract: Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either hardly alter models'internal knowledge or h

LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization

AgentsDGX agent

arXiv:2607.20503v1 Announce Type: new Abstract: We present and evaluate LeanFlow, an LLM agent system specialized for translating mathematical papers into buildable Lean projects. Recent verifier-in-t

Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc

Model ReleasesDGX agent

arXiv:2607.20456v1 Announce Type: cross Abstract: Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such as MiniZinc

Leveraging Biokinetic Knowledge Priors for Data-Scarce Bioprocess Modeling

TutorialsDGX agent

arXiv:2607.20539v1 Announce Type: cross Abstract: While deep learning has accelerated drug discovery, its impact on biomanufacturing has been considerably more limited. The reason is data scarcity. Bi

LinearARD: Linear-Memory Attention Distillation for RoPE Restoration

ResearchDGX agent

arXiv:2604.00004v2 Announce Type: replace-cross Abstract: The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by lightweight Con

LLM-INSTRUCT at UZH Shared Task 2026: Constraint-Aware Retrieval and Selective Debate for Paragraph-Level Argument Mining

AgentsDGX agent

arXiv:2607.20430v1 Announce Type: cross Abstract: We present LLM-INSTRUCT, the winning system for the UZH Shared Task at ArgMining 2026 on paragraph-level argument mining in UN and UNESCO resolutions.

Logic Programming Semantics for Causal Processes

ResearchDGX agent

arXiv:2607.21233v1 Announce Type: new Abstract: Motivated by challenging modelling issues in the life sciences, we investigate the relationship between logic programming semantics and the eventual sta

Logical Regression for Planning with Axioms

ResearchDGX agent

arXiv:2607.21414v1 Announce Type: new Abstract: In automated planning, logical regression is an operation that returns the most general condition necessary for an action to achieve a particular formul

Loss-Complexity Landscape and Model Structure Functions

ResearchDGX agent

arXiv:2507.13543v5 Announce Type: replace-cross Abstract: We develop a framework for dualizing the Kolmogorov structure function h_x(alpha), which then allows using computable complexity proxies. We e

M^3-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data

TutorialsDGX agent

arXiv:2607.21343v1 Announce Type: cross Abstract: Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for comprehensive di

Making Open-Source Text LLM Watermarks Durable Against Merging

TutorialsDGX agent

arXiv:2607.20435v1 Announce Type: cross Abstract: Open-source LLMs (OSMs)arereaching near state-of-the-art performance, prompting prior works to trace the text they generate by embedding text watermar

Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts

ResearchDGX agent

arXiv:2607.20462v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated output

Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy

Model ReleasesDGX agent

arXiv:2607.21372v1 Announce Type: cross Abstract: Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. While positivity guarantees

MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs

SafetyDGX agent

arXiv:2508.05502v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally 'thin' descripti

MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference

AgentsDGX agent

arXiv:2607.20507v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for program-aided reasoning, agentic decision making, and structured task execution, but these applic

Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation

Local AiDGX agent

arXiv:2512.07540v4 Announce Type: replace-cross Abstract: Error Span Detection (ESD) extends automatic machine translation (MT) evaluation by localizing translation errors and labeling their severity.

MIRROR: Learning from the Other View for Multi-Modal Reasoning

ResearchDGX agent

arXiv:2607.21552v1 Announce Type: new Abstract: Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on ge

MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation

AgentsDGX agent

arXiv:2607.20501v1 Announce Type: new Abstract: Despite rapid progress in LLM-based code generation, writing correct and performant kernels for hardware accelerators remains a key bottleneck in scalin

Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing

Model ReleasesDGX agent

arXiv:2607.20433v1 Announce Type: cross Abstract: While language models remain frozen at their training state, the world evolves continuously. Knowledge editing has emerged as a key alternative to ful

Monkey King Bang: A Unified Scientific Multimodal Foundation Model

ResearchDGX agent

arXiv:2607.20557v1 Announce Type: cross Abstract: Scientific discovery is increasingly shifting from isolated disciplines to multi-domain reasoning, and AI for science faces a similar transition. Exis

More Is Not More: What Matters for Diversity in LLM Opinions?

ResearchDGX agent

arXiv:2607.20429v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, an

MSBraM: A Multi-scale Self-supervised Brain Foundation Model for Hierarchical EEG Dynamics Learning

Local AiDGX agent

arXiv:2607.21402v1 Announce Type: new Abstract: Self-supervised foundation models have recently shown strong potential for electroencephalogram (EEG)-based analysis. However, existing approaches strug

Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning

ResearchDGX agent

arXiv:2607.21290v1 Announce Type: cross Abstract: Multi-task learning (MTL) is a promising approach for prediction tasks derived from video game state data, as modern game telemetry provides multiple

Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation

Model ReleasesDGX agent

arXiv:2607.20908v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for optimized code

Multilevel Graph Wavelet Compressed Sensing with Scale-Aware Neural Recovery

ResearchDGX agent

arXiv:2607.20857v1 Announce Type: cross Abstract: Scientific machine learning methods such as neural operators and physics-informed neural networks have advanced engineering applications and inverse p

Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems

ResearchDGX agent

arXiv:2602.08792v2 Announce Type: replace-cross Abstract: The pantograph-catenary interface is essential for ensuring uninterrupted and reliable power delivery in electrified rail systems. However, el

Multimodal Pretraining for Generalizable EEG Representation Learning

Model ReleasesDGX agent

arXiv:2607.21384v1 Announce Type: new Abstract: Electroencephalography (EEG) models used for epilepsy are often limited to specific datasets and tasks. This limited approach can make it challenging to

Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory

Local AiDGX agent

arXiv:2607.21000v1 Announce Type: new Abstract: Long-sequence memory tracking places two opposing demands on a recurrent state: near-lossless retention of stored bindings over long horizons, and activ

NeuraLSP: A Neural Spectral Preconditioner for Accelerating PDE Solvers

TutorialsDGX agent

arXiv:2601.20174v3 Announce Type: replace-cross Abstract: Solving large-scale sparse linear systems originating from partial differential equations (PDEs) is a fundamental topic in high-performance sc

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

HardwareDGX agent

arXiv:2607.20709v1 Announce Type: new Abstract: Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agen

On the Granularity of Causal Effect Identifiability

ResearchDGX agent

arXiv:2510.16703v3 Announce Type: replace-cross Abstract: The classical notion of causal effect identifiability is defined in terms of treatment and outcome variables. In this paper, we consider the i

One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies

Model ReleasesDGX agent

arXiv:2607.21143v1 Announce Type: cross Abstract: Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to a

OpenForgeRL: Train Harness-native Agents in Any Environment

Model ReleasesDGX agent

arXiv:2607.21557v1 Announce Type: new Abstract: Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to e

Operational Identity: A Finite Audit of Declared and Implemented Rules of Sameness

SafetyDGX agent

arXiv:2607.20729v1 Announce Type: cross Abstract: A record system declares when two records refer to the same entity, occurrence, scope, or rule. Its disclosed implementation mechanisms induce a corre

OPOD: On-Policy Omni Distillation

SafetyDGX agent

arXiv:2607.20918v1 Announce Type: new Abstract: Omni-modal models can handle text, images, and audio in one system, but improving all of these abilities together remains difficult. Training a single m

Optimizing Hypergraph-Based RAG: Toward Better Fact Extraction and Chunk Retrieval

ResearchDGX agent

arXiv:2607.20506v1 Announce Type: new Abstract: GraphRAG enables deeper reasoning by structuring knowledge as graphs but struggles with n-ary facts. HyperGraphRAG uses hypergraphs for richer semantics

OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining

AgentsDGX agent

arXiv:2607.20486v1 Announce Type: new Abstract: Designing optimizers for modern deep learning remains a challenging scientific problem, requiring the joint consideration of optimization geometry, stat

pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development

Model ReleasesDGX agent

arXiv:2607.21268v1 Announce Type: cross Abstract: In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable co

PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2607.21419v1 Announce Type: new Abstract: In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories and limiting

PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided Editing

ResearchDGX agent

arXiv:2607.21318v1 Announce Type: cross Abstract: Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by the source

PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

Model ReleasesDGX agent

arXiv:2607.20482v1 Announce Type: new Abstract: Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspeci

PhantomFill: When the Form Demands an Answer, Language Models Invent One

Model ReleasesDGX agent

arXiv:2607.20492v1 Announce Type: cross Abstract: Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself

Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin

SafetyDGX agent

arXiv:2607.21332v1 Announce Type: cross Abstract: Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource language va

PILD: Physics-Informed Learning via Diffusion

SafetyDGX agent

arXiv:2601.21284v2 Announce Type: replace-cross Abstract: Diffusion models have emerged as powerful generative tools for modeling complex data distributions, yet their purely data-driven nature limits

PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs

ResearchDGX agent

arXiv:2607.20470v1 Announce Type: new Abstract: Enhancing the task-specific capabilities of Large Language Models (LLMs) primarily requires substantial instruction-tuning datasets. However, the sheer

Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development

ResearchDGX agent

arXiv:2607.20532v1 Announce Type: cross Abstract: Many modern AI systems are designed to operate under diverse, open-ended, use-cases. To help generalize deployed systems, many deployed-system mainten

Post-Hoc Reasoning in Chain of Thought: Decoding and Steering Pre-Committed Answers

ResearchDGX agent

arXiv:2603.01437v2 Announce Type: replace Abstract: As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising t

Preference Tuning as Spectral Update Reorganization

Model ReleasesDGX agent

arXiv:2607.20438v1 Announce Type: cross Abstract: Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains largely opa

Probabilistic Residual Learning for Online Recommendations

ResearchDGX agent

arXiv:2607.20863v1 Announce Type: cross Abstract: Modern recommender systems are typically based on deep learning (DL) models, where a dense encoder learns representations of users and items. As a res

Profiling Lightweight Large Language Models

Model ReleasesDGX agent

arXiv:2607.20806v1 Announce Type: new Abstract: Lightweight large language models (LLMs) are increasingly being deployed locally on personal computers and are expected to play a growing role in resour

PromptPack: Scaling LLM Annotation Agents for Online Recommendation

Model ReleasesDGX agent

arXiv:2607.20528v1 Announce Type: new Abstract: Online recommendation platforms increasingly use Large Language Models (LLMs) to extract structured features from ad creatives. While deploying a single

RE-AD: Real-Time Requirement Adherence for Data Labeling

Model ReleasesDGX agent

arXiv:2607.20455v1 Announce Type: cross Abstract: Human-annotated data remains fundamental to training frontier Large Language Models (LLMs). However, crowd-sourced annotations often suffer from quali

RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring

ApplicationsDGX agent

arXiv:2607.20628v1 Announce Type: cross Abstract: Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the scarcity of realistic training data, yet

← Previous
1…5657585960…354
Next →