AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
3 Jul 2026

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Model ReleasesDGX agent

arXiv:2607.02440v1 Announce Type: new Abstract: Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a

Expander Sparse Autoencoders: Parameter-Efficient Dictionaries for Mechanistic Interpretability

Model ReleasesDGX agent

arXiv:2607.01799v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) decompose internal activations of neural networks into sparse linear combinations of learned features by fitting an overcom

ExPerT: Personalizing LLM Responses to Users' Domain Expertise via Query-Wise Semantic and Keystroke Behavioral Cues

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.01242v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used by end users, yet existing personalization methods relying on static profiles or text-only signals

Exploring Large Language Models for Access Control Policy Synthesis and Summarization

SafetyDGX agent

arXiv:2510.20692v2 Announce Type: replace-cross Abstract: Cloud computing is ubiquitous, with a growing number of services being hosted on the cloud every day. Typical cloud compute systems allow admi

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

Model ReleasesDGX agent

arXiv:2607.02396v1 Announce Type: new Abstract: Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behav

From Experiments to Expertise: Scientific Knowledge Consolidation for AI-Driven Computational Physics

ResearchDGX agent

arXiv:2603.13191v2 Announce Type: replace-cross Abstract: While large language models (LLMs) have transformed AI agents into proficient executors of computational materials science, performing a hundr

Full Bayesian Reinforcement Learning via LF-IBIS

SafetyDGX agent

arXiv:2607.01741v1 Announce Type: cross Abstract: Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an environmen

Fully Unsupervised Detection of Physical Contacts on Subsea Cables via State-of-Polarization Monitoring

ResearchDGX agent

arXiv:2607.01484v1 Announce Type: cross Abstract: We present a fully unsupervised Fast-Slow DSVDD detector for continuous State-of-Polarization monitoring on a deployed subsea cable. Trained without e

G-RRM: Guiding Symbolic Solvers with Recurrent Reasoning Models

TutorialsDGX agent

arXiv:2607.02491v1 Announce Type: new Abstract: In this work, we focus on SE-RRMs, a symbol-equivariant instantiation of RRMs that exhibits improved extrapolation to larger problem sizes. We propose a

GAP-GDRNet: Geometry-Aware Monocular Visual Pose Sensing on a Single-Target Synthetic Spacecraft Dataset

Local AiDGX agent

arXiv:2607.02360v1 Announce Type: cross Abstract: Monocular relative pose sensing is a central perception problem in non-cooperative rendezvous and on-orbit servicing. In spacecraft images, however, w

Generalization in offline RL: The structure is more important than the amount of pessimism

SafetyDGX agent

arXiv:2607.02288v1 Announce Type: cross Abstract: While pessimism counteracts overestimation bias in offline reinforcement learning (RL), being overly conservative has been associated with hindering c

Generative AI and Federated Learning for Intrusion Detection Systems: A Survey

Model ReleasesDGX agent

arXiv:2607.01305v1 Announce Type: cross Abstract: Intrusion Detection Systems (IDSs) are essential for monitoring network traffic and identifying malicious activities in modern cyber-physical, Interne

Generic Expert Coverage for Pruning SparseMixture-of-Experts Language Models

Model ReleasesDGX agent

arXiv:2607.01710v1 Announce Type: new Abstract: Sparsely activated Mixture-of-Experts (MoE) language models contain substantial structured redundancy among routed experts, but pruning them without dow

GPUAlert: A Zero-Instrumentation Process-Boundary Monitor for Diagnosing GPU Training-Job Failures

HardwareDGX agent

arXiv:2607.01409v1 Announce Type: cross Abstract: GPU training jobs fail often, roughly two in five on large production clusters, yet the operator typically learns of a failure only by reconnecting ho

Gravity-Awareness: Deep Learning Models and LLM Simulation of Human Awareness in Altered Gravity

Model ReleasesDGX agent

arXiv:2511.05536v2 Announce Type: replace-cross Abstract: Earth s gravity fundamentally shapes human behaviour. The brain encodes this force as an internal model of gravity, enabling the prediction an

Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics

AgentsDGX agent

arXiv:2607.02329v1 Announce Type: new Abstract: Autonomous-research agents have demonstrated end-to-end LLM automation in machine-learning sandboxes where execution provides calibration. Frontier phys

Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting

AgentsDGX agent

arXiv:2607.01457v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly applied to resume optimization for applicant tracking systems, introducing hallucination failures distin

GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation

Model ReleasesDGX agent

arXiv:2606.22737v2 Announce Type: replace Abstract: Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic tes

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies

SafetyDGX agent

arXiv:2607.02092v1 Announce Type: cross Abstract: Flow-matching vision-language-action policies generate robot action chunks through an iterative transport process, creating an opportunity for test-ti

HAL: Inducing Human-likeness in LLMs with Alignment

SafetyDGX agent

arXiv:2601.02813v3 Announce Type: replace Abstract: Aligning language models to qualitative behavioral traits, such as human-likeness, remains difficult because they are hard to define, measure, and o

Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems

SafetyDGX agent

arXiv:2607.02376v1 Announce Type: new Abstract: Recent advances in agentic AI are producing increasingly complex autonomous systems that integrate large language models, world models, optimization eng

HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

Model ReleasesDGX agent

arXiv:2606.14249v2 Announce Type: replace Abstract: AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model obs

Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map

Model ReleasesDGX agent

arXiv:2607.01854v1 Announce Type: cross Abstract: Can a platform tell, before deployment, whether an open-weight checkpoint has had its refusal mechanism stripped? Runtime guards cannot: they score ge

Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

ApplicationsDGX agent

arXiv:2607.01590v1 Announce Type: new Abstract: Developing high-performance kernels for Neural Processing Units (NPUs) is a critical industry bottleneck, requiring developers to manually navigate impl

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures

Model ReleasesDGX agent

arXiv:2607.02266v1 Announce Type: cross Abstract: Most data-mixing methods assume the corpus has already been partitioned into groups, and the choice of those groups determines what a mixer can expres

Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails

ResearchDGX agent

arXiv:2607.02020v1 Announce Type: new Abstract: Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether ol

Horizon-Uniform Sensitivity Certificates for Finite-Horizon Pontryagin Systems

ResearchDGX agent

arXiv:2606.17762v2 Announce Type: replace-cross Abstract: Finite-horizon optimal-control computations repeatedly solve two-point Pontryagin boundary value problems whose conditioning can deteriorate a

How Indian Dermatologists are Utilizing Artificial Intelligence for Clinical Practice and Workflow Management: A Nationwide Survey with a Special Focus on atopic dermatitis

ResearchDGX agent

arXiv:2607.01252v1 Announce Type: cross Abstract: Background: Dermatology AI has mainly focused on image-based diagnosis, while chronic disease workflows have received less attention. We surveyed Indi

How Should Transformers Encode Numeric Values in Electronic Health Records?

ApplicationsDGX agent

arXiv:2607.01391v1 Announce Type: cross Abstract: How do we encode numeric values in transformer-based sequence processing, particularly in electronic health record (EHR) data? We systematically compa

Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting

Model ReleasesDGX agent

arXiv:2607.02467v1 Announce Type: cross Abstract: Whether pairing people with AI helps or hurts is usually reported as a single average effect. Using a real-money prediction market (Polymarket) as an

Hybrid Diffusion Transformer for Instruction-Guided Audio Editing via Rectified Flow

Local AiDGX agent

arXiv:2606.20101v2 Announce Type: replace-cross Abstract: Audio editing aims to modify specific content in an existing audio clip according to a natural language instruction while preserving the remai

InduceKV: Fixed-Footprint Continual Adaptation of Multimodal LLMs via Inducing KV Memories

Model ReleasesDGX agent

arXiv:2607.02010v1 Announce Type: new Abstract: Multimodal large language models must adapt to evolving tasks and domains, yet continual improvement under bounded deployment footprint remains difficul

IntentTune: Using user demand and personalization to resolve 'unknown' query intents for e-commerce search

ApplicationsDGX agent

arXiv:2607.01530v1 Announce Type: cross Abstract: Understanding user intent is fundamental to delivering relevant search results in e-commerce. However, substantial fraction of real-world queries are

Interpreting Global Perturbation Robustness of Image Models using Axiomatic Spectral Importance Decomposition

ResearchDGX agent

arXiv:2408.01139v4 Announce Type: replace Abstract: Perturbation robustness evaluates the vulnerabilities of models, arising from a variety of perturbations, such as data corruptions and adversarial a

Introduction to Transformers: an NLP Perspective

ResearchDGX agent

arXiv:2311.17633v2 Announce Type: replace-cross Abstract: Transformers have dominated empirical machine learning models of natural language processing. In this paper, we introduce basic concepts of Tr

IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs

Model ReleasesDGX agent

arXiv:2607.01431v1 Announce Type: cross Abstract: We introduce ISOSCI, a benchmark of isomorphic cross-domain science problem pairs that separates reasoning ability from domain knowledge retrieval in

Janus: a Playground for User-Involved Agentic Permission Management

AgentsDGX agent

arXiv:2607.01510v1 Announce Type: new Abstract: AI agents that autonomously execute tool calls on a user's behalf raise pressing questions about permission management: what role could users play, and

Kara: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression

SafetyDGX agent

arXiv:2607.01237v1 Announce Type: cross Abstract: Reasoning language models often generate long chain-of-thought (CoT), which accumulates a massive KV cache during the decoding phase and incurs high d

kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail

SafetyDGX agent

arXiv:2607.02072v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in domains requiring guardrails to detect unsafe, off-topic, or adversarial prompts. Existing g

LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

Model ReleasesDGX agent

arXiv:2607.02513v1 Announce Type: cross Abstract: LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal met

Learning 3D-Gaussian Simulators from RGB Videos

TutorialsDGX agent

arXiv:2503.24009v3 Announce Type: replace-cross Abstract: Realistic simulation is critical for applications ranging from robotics to animation. Learned simulators have emerged as a possibility to capt

Learning-based Multi-agent Race Strategies in Formula 1

SafetyDGX agent

arXiv:2602.23056v2 Announce Type: replace Abstract: In Formula 1, race strategies are adapted according to evolving race conditions and competitors' actions. This paper proposes a reinforcement learni

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Model ReleasesDGX agent

arXiv:2607.02466v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions,

LEFT: Learnable Fusion of Tri-view Tokens for Unsupervised Time Series Anomaly Detection

TutorialsDGX agent

arXiv:2602.08638v2 Announce Type: replace-cross Abstract: As a fundamental data mining task, unsupervised time series anomaly detection (TSAD) aims to build a model for identifying abnormal timestamps

Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens

Model ReleasesDGX agent

arXiv:2507.02964v2 Announce Type: replace-cross Abstract: The increasing scale of AI workloads demands High-Performance Computing (HPC) infrastructure and training methodologies that are both scalable

Leveraging Metamemory Agent for Enhanced Data-Free Code Generation in Large Language Models

AgentsDGX agent

arXiv:2501.07892v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong performance in automated code generation, with few-shot prompting widely used for its simplicit

Lightweight Safe Reinforcement Learning for End-to-End UAV Navigation

SafetyDGX agent

arXiv:2607.01794v1 Announce Type: cross Abstract: With the rapid development of autonomous aerial systems, Unmanned Aerial Vehicles (UAVs) are increasingly deployed in applications such as inspection,

LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability

Model ReleasesDGX agent

arXiv:2607.01247v1 Announce Type: cross Abstract: Open-ended mathematics exams are valuable because they assess reasoning, proof construction, algorithmic thinking, and communication of intermediate s

Locality-Aware Continual Unlearning for Diffusion Models

Model ReleasesDGX agent

arXiv:2512.02657v2 Announce Type: replace-cross Abstract: Real-world deployment of text-to-image diffusion models requires continual concept removal as new privacy, copyright, or safety obligations ar

Low-Latency Task-Oriented Image Transmission with Opportunistic Spectrum Access

ResearchDGX agent

arXiv:2607.01921v1 Announce Type: cross Abstract: Communication systems designed for reliable data reconstruction, rather than task-oriented communication, typically rely on separate source and channe

MAGIK: Mapping to Analogous Goals via Imagination-enabled Knowledge Transfer

SafetyDGX agent

arXiv:2506.01623v4 Announce Type: replace Abstract: Humans excel at analogical reasoning - applying knowledge from one task to a related one with minimal relearning. In contrast, reinforcement learnin

Mapping Text to Multiplex Graph: Prompt Compression as Levy Walk-Guided Graph Pruning

Local AiDGX agent

arXiv:2607.01241v1 Announce Type: cross Abstract: Existing prompt compression methods treat text as flat token sequences, failing to capture the distributed nature of important information, which is o

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction

Model ReleasesDGX agent

arXiv:2607.01764v1 Announce Type: new Abstract: Repository-level vulnerability reproduction is a demanding software engineering (SE) task: an agent must inspect a codebase, infer the input grammar tha

Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders

ResearchDGX agent

arXiv:2607.01336v1 Announce Type: cross Abstract: Neural Quantum States (NQS) are a remarkably expressive class of variational ansatze for quantum many-body wavefunctions, yet little is understood abo

MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation

Model ReleasesDGX agent

arXiv:2508.16674v2 Announce Type: replace-cross Abstract: Medical report understanding from real-world document images is essential for generating patient-facing explanations and enabling structured i

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

Model ReleasesDGX agent

arXiv:2607.01751v1 Announce Type: cross Abstract: Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right ti

Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity

SafetyDGX agent

arXiv:2602.03315v2 Announce Type: replace Abstract: Agent memory systems must accommodate continuously growing information while supporting efficient, context-aware retrieval for downstream tasks. Abs

Meta-Benchmarks for Financial-Services LLM Evaluation

Model ReleasesDGX agent

arXiv:2607.01740v1 Announce Type: new Abstract: Public LLM leaderboards optimise for global average performance and do not capture the specific cognitive demands of financial-services work: a model th

MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2506.09105v3 Announce Type: replace-cross Abstract: We present MetaTT, a Tensor Train (TT) adapter framework for fine-tuning of pre-trained transformers. MetaTT enables flexible and parameter-ef

Mirror Illusion Art

SafetyDGX agent

arXiv:2607.02015v1 Announce Type: cross Abstract: Mirror Illusion Art is a novel reflection-conditioned 3D illusion where one object yields two target appearances (front and mirror). The task is formu

← Previous
1…9596979899…358
Next →