AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
28 Jul 2026

AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use

Model ReleasesDGX agent

arXiv:2505.12650v2 Announce Type: replace-cross Abstract: Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can explain

AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models

AgentsDGX agent

arXiv:2603.28963v2 Announce Type: replace-cross Abstract: Simulation with realistic traffic agents is essential for validating autonomous driving systems. Existing data-driven simulators learn agent b

Bayesian Repetition Penalty: A Principled Adjacent-Conditional Framework for Reversing Attention Collapse in Autoregressive Language Models

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2607.22694v1 Announce Type: new Abstract: Attention collapse in autoregressive language models -- manifested as repetitive token loops where the model becomes trapped in self-reinforcing attract

BettiSplit: Topology-Guided Privacy-Aware Split Learning Against Feature Inversion and Gradient Leakage

ApplicationsDGX agent

arXiv:2607.24556v1 Announce Type: cross Abstract: Split learning enables collaborative model training by partitioning neural networks across clients and servers. However, improper split placement can

Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS

Model ReleasesDGX agent

arXiv:2607.22657v1 Announce Type: cross Abstract: Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing

Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining

SafetyDGX agent

arXiv:2607.23175v1 Announce Type: cross Abstract: Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present

Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls

ResearchDGX agent

arXiv:2607.24343v1 Announce Type: cross Abstract: Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an email body bu

Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models

HardwareDGX agent

arXiv:2607.22663v1 Announce Type: new Abstract: Block diffusion has emerged as the dominant paradigm for scaling discrete diffusion language models (dLLMs), because decoding text in fixed-size blocks

Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.22996v1 Announce Type: cross Abstract: Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the opening turn inst

Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents

Model ReleasesDGX agent

arXiv:2607.22689v1 Announce Type: new Abstract: Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions th

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

Model ReleasesDGX agent

arXiv:2607.22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuni

Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks

SafetyDGX agent

arXiv:2410.02596v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) are a novel class of generative models designed to sample from unnormalized distributions and have found

Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents

AgentsDGX agent

arXiv:2607.15263v3 Announce Type: replace-cross Abstract: Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, e

Blood Pressure Estimation from PPG: A Comparative Study of Direct and ECG-Mediated Deep Learning Pipelines

ApplicationsDGX agent

arXiv:2607.23406v1 Announce Type: cross Abstract: Continuous cuffless blood pressure (BP) monitoring is essential for connected health systems and wearable devices, enabling early detection, longitudi

BoneAgeTW2: Automated Skeletal Maturation Assessment via the Tanner-Whitehouse 2 Method, Deep Learning, and Clinical Report Generation with Distribution Curves

Local AiDGX agent

arXiv:2607.23224v1 Announce Type: cross Abstract: We present BoneAgeTW2, the first fully open-source system to automate the complete Tanner-Whitehouse 2 (TW2) clinical protocol for skeletal maturity a

Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational Excellence

AgentsDGX agent

arXiv:2607.22948v1 Announce Type: cross Abstract: The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 initiative and

CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

Model ReleasesDGX agent

arXiv:2607.23159v1 Announce Type: new Abstract: Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although most are discard

CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding

Local AiDGX agent

arXiv:2607.24582v1 Announce Type: cross Abstract: Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference p

CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants

Model ReleasesDGX agent

arXiv:2607.22635v1 Announce Type: new Abstract: Target-oriented dialogue systems have demonstrated strong capabilities in completing user goals through interactive conversations. However, existing stu

CALMRec: Causally Aligned Language Memory for Long-Horizon Recommendation

Model ReleasesDGX agent

arXiv:2607.23647v1 Announce Type: cross Abstract: Large language models (LLMs) can summarize heterogeneous user evidence in natural language, but current LLM recommenders often collapse enduring prefe

Capacity-Aware Deep Learning for Generalizable Traffic Volume Estimation Across Links and Cities

ResearchDGX agent

arXiv:2607.24056v1 Announce Type: cross Abstract: Network-wide traffic volume estimation typically relies on propagating measurements from fixed sensors, making performance highly dependent on sensor

CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics

ResearchDGX agent

arXiv:2607.23258v1 Announce Type: new Abstract: Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central question remains u

CausAdv: A Causal-based Framework for Detecting Adversarial Examples

ResearchDGX agent

arXiv:2411.00839v4 Announce Type: replace-cross Abstract: Deep learning has led to tremendous success in computer vision, largely due to Convolutional Neural Networks (CNNs). However, CNNs have been s

Characterisation of Density-based FM generation methods in the context of Information Fusion

ResearchDGX agent

arXiv:2607.23243v1 Announce Type: new Abstract: Fuzzy Integral (FI) based aggregation provides a powerful mechanism for nuanced aggregation, for example, in ensemble approaches or decision-level fusio

Chart Deception in Vision-Language Models: From Vulnerability to Mitigation

Model ReleasesDGX agent

arXiv:2607.22600v1 Announce Type: new Abstract: Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or inverted axe

Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models

Model ReleasesDGX agent

arXiv:2607.22771v1 Announce Type: cross Abstract: Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a search over m

Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework

Model ReleasesDGX agent

arXiv:2607.23507v1 Announce Type: cross Abstract: Choosing the right text embedding model is one of the most consequential -- and most frequently under-examined -- decisions in building a retrieval or

CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process

HardwareDGX agent

arXiv:2607.22624v1 Announce Type: new Abstract: Recently, there have been several works in the Text-to-SQL domain that utilize Small Language Models (SLMs) for training. These approaches achieve perfo

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

Model ReleasesDGX agent

arXiv:2607.24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundam

Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs

Model ReleasesDGX agent

arXiv:2607.24371v1 Announce Type: cross Abstract: Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic codin

cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs

Model ReleasesDGX agent

arXiv:2607.22577v1 Announce Type: new Abstract: Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for

Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents

Local AiDGX agent

arXiv:2607.22688v1 Announce Type: new Abstract: Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how research traj

Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification

ApplicationsDGX agent

arXiv:2607.24683v1 Announce Type: cross Abstract: Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real-world scen

CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback

AgentsDGX agent

arXiv:2507.22080v2 Announce Type: replace-cross Abstract: Acquiring high-quality instruction-code pairs is essential for training Large Language Models for code generation. While automated synthesis h

CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases

AgentsDGX agent

arXiv:2408.03910v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories. Thi

Codifying the Judge: Scalable Evaluation via Program Distillation

ResearchDGX agent

arXiv:2607.22561v1 Announce Type: new Abstract: LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations

Coherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure

AgentsDGX agent

arXiv:2603.28371v2 Announce Type: replace-cross Abstract: When an agent can articulate why something works, we typically take this as evidence of genuine understanding. This presupposes that effective

Commitment To Cooperation With Self-Negotiated Contracts

AgentsDGX agent

arXiv:2607.22750v1 Announce Type: new Abstract: As AI agents operate with increasing autonomy in a multi-agent world, they will need to learn to cooperate with other agents and with humans to generate

Comparing Optimization Models for Radiotherapy Scheduling

ResearchDGX agent

arXiv:2607.22539v1 Announce Type: cross Abstract: The Radiotherapy Scheduling Problem (RTSP) involves determining an optimal schedule for patients undergoing radiation treatments, a task that has a ma

Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization

Model ReleasesDGX agent

arXiv:2607.23089v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled automated kernel generation and optimization, but most existing approaches rely on surface

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV

AgentsDGX agent

arXiv:2607.23693v1 Announce Type: new Abstract: Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries and drops the rest. Eviction and epis

Concept-based Visual Counterfactual Explanations with Diffusion Models

SafetyDGX agent

arXiv:2607.22544v1 Announce Type: new Abstract: Visual counterfactual explanations aim to answer 'what minimal change to this image would flip the model's prediction?', and are increasingly important

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

Model ReleasesDGX agent

arXiv:2607.23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional

ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control

Model ReleasesDGX agent

arXiv:2607.22962v1 Announce Type: new Abstract: LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A hallucinated

Constraint-Bound Agnostic Bayesian Optimization: One Model for All Thresholds

Model ReleasesDGX agent

arXiv:2607.23448v1 Announce Type: cross Abstract: Expensive constrained optimization problems in real-world industry design often involve constraint thresholds that are difficult to determine in advan

Context-Aware Concept Distillation for Trustworthy Flood Prediction

Local AiDGX agent

arXiv:2607.23237v1 Announce Type: cross Abstract: Effective flood risk management relies on accurate forecasting, yet the 'black box' nature of stateof-the-art Deep Learning models creates a barrier t

Continual Knowledge Consolidation LORA for Domain Incremental Learning

Model ReleasesDGX agent

arXiv:2510.16077v2 Announce Type: replace-cross Abstract: Domain Incremental Learning (DIL) is a sub-branch of continual learning that aims to address the never-ending arrival of new domains without c

Controlling Embedding Spaces with Text-Conditioned Transformations

TutorialsDGX agent

arXiv:2607.22919v1 Announce Type: cross Abstract: Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classific

CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning

SafetyDGX agent

arXiv:2505.18334v2 Announce Type: replace-cross Abstract: Past work has demonstrated that autonomous vehicles can drive more safely if they communicate with each other. However, this communication is

Coordinated Networking for On-Device Agent-Augmented Real-Time Communication

Model ReleasesDGX agent

arXiv:2607.22854v1 Announce Type: new Abstract: AI agents are enabling a new paradigm of agent-augmented real-time communication (RTC), where humans focus on high-level collaboration, while agents aut

Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features

Model ReleasesDGX agent

arXiv:2607.22739v1 Announce Type: cross Abstract: We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcement learnin

CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents

AgentsDGX agent

arXiv:2607.22711v1 Announce Type: cross Abstract: LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. Howeve

Cost-Aware Recovery-Pathway Identification and Bayesian Optimization for Autonomous Materials Discovery

Model ReleasesDGX agent

arXiv:2607.23896v1 Announce Type: new Abstract: Autonomous laboratories automate experimental execution, but a campaign must also decide which recovery pathway merits optimization. We formulate this a

CRAFT: Learn the Schema, Execute the Plan

SafetyDGX agent

arXiv:2607.22642v1 Announce Type: new Abstract: Enterprise coding agents translate natural-language analytical requests into executable code over proprietary APIs, schemas, and metric definitions. Yet

Creative Integration: A Decidable Criterion of Creativity

ResearchDGX agent

arXiv:2606.13977v1 Announce Type: cross Abstract: 'Integrative' solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes the world ch

CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference

HardwareDGX agent

arXiv:2511.21702v2 Announce Type: replace-cross Abstract: Large language models face significant computational bottlenecks during inference due to the expensive output layer computation over large voc

CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data

ResearchDGX agent

arXiv:2607.22662v1 Announce Type: new Abstract: Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have significantly

D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models

Model ReleasesDGX agent

arXiv:2607.24586v1 Announce Type: cross Abstract: Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to b

D3O: Dynamic Distribution Distillation for Ordinal Regression

SafetyDGX agent

arXiv:2607.23575v1 Announce Type: cross Abstract: Ordinal regression is widely used in scenarios where labels are discrete yet inherently ordered. In practice, however, ordinal labels are often obtain

DAMamba-UNet3D: A Parameter-Efficient Mamba State Space U-Net with Dynamic Adaptive Scan for 3D Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2607.22718v1 Announce Type: cross Abstract: We propose parameter-efficient SSM-based U-Net architectures for 3D medical image segmentation. Convolutional U-Nets afford O(n) local mixing per laye

← Previous
1…4647484950…354
Next →