AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
6 Jun 2026

Can LLMs Write Correct TLA+ Specifications? Evaluating Natural-Language-to-TLA+ Generation

Model ReleasesDGX agent

arXiv:2606.05792v1 Announce Type: new Abstract: TLA+ has supported industrial verification at companies such as Amazon and Microsoft, yet writing correct TLA+ specifications from natural language stil

CangLing-KnowFlow: A Unified Knowledge-and-Flow-fused Agent for Comprehensive Remote Sensing Applications

Model ReleasesDGX agent

arXiv:2512.15231v3 Announce Type: replace Abstract: The automated and intelligent processing of massive remote sensing (RS) datasets is critical in Earth observation (EO). Existing automated systems a

Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2606.05966v1 Announce Type: cross Abstract: Understanding and reasoning about the physical world is the foundation of intelligent behavior, yet state-of-the-art vision-language models (VLMs) sti

CausalPOI: Spatio-Temporal Graph-Based Causal Modeling for Cold-Start POI Check-in Forecasting

ApplicationsDGX agent

arXiv:2606.05413v1 Announce Type: cross Abstract: As urban environments continue to evolve rapidly, accurately modeling the dynamic behaviour of Points of Interest is essential for supporting data-dri

Class-Specific Branch Attention for Mitigating Gradient Interference under Class Imbalance

SafetyDGX agent

arXiv:2606.05740v1 Announce Type: new Abstract: Deep neural networks trained under severe class imbalance often exhibit degraded performance, typically attributed to statistical bias. In this work, we

Closing the Loop on Latent Reasoning via Test-Time Reconstruction

Model ReleasesDGX agent

arXiv:2606.06252v1 Announce Type: new Abstract: Recent work moves intermediate reasoning from natural-language traces into latent or cache-level representations to reduce token overhead and avoid a di

CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model

Model ReleasesDGX agent

arXiv:2606.06099v1 Announce Type: new Abstract: Whether Large Language Models (LLMs) exhibit covert psychological manipulation in complex human-AI interactions has garnered increasing safety concerns.

Cognitive Threat Intelligence and Explainable Federated Security Analytics for distributed Infrastructure Systems

Local AiDGX agent

arXiv:2606.05701v1 Announce Type: cross Abstract: The increasing adoption of distributed infrastructure systems, cloud computing, Internet of Things (IoT) technologies, and edge-based architectures ha

Compositional Boundaries for Density Fusion

Local AiDGX agent

arXiv:2606.05871v1 Announce Type: cross Abstract: Distributed uncertainty-management systems often combine local probabilistic models along aggregation trees chosen by communication, privacy, or sched

Comprehensive and Reliable Feature Attribution for Diverse Modalities and Models via Frequency-Domain Insights

SafetyDGX agent

arXiv:2411.18343v3 Announce Type: replace-cross Abstract: Personalized Federal learning(PFL) allows clients to cooperatively train a personalized model without disclosing their private dataset. Howeve

Conformal Risk-Averse Decision Making with Action Conditional Guarantee

SafetyDGX agent

arXiv:2606.05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety g

Consistency Training Along the Transformer Stack

SafetyDGX agent

arXiv:2606.05817v1 Announce Type: cross Abstract: Consistency training encourages models to behave similarly across different contexts, and has shown promise for reducing misalignment. We broaden the

Critic-Guided Heterogeneous Multi-Agent Reasoning for Reliable Mathematical Problem Solving

Model ReleasesDGX agent

arXiv:2606.05704v1 Announce Type: new Abstract: Recent Large Language Models (LLMs) have shown impressive reasoning abilities; but they are still susceptible to hallucinations, intermediate reasoning

Cross-Epoch Adaptive Rollout Optimization for RL Post-Training

Model ReleasesDGX agent

arXiv:2606.05606v1 Announce Type: cross Abstract: LLM post-training often relies on reinforcement learning methods that sample multiple rollouts per prompt, yet most existing approaches use a fixed ro

CTIConnect: A Benchmark for Retrieval-Augmented LLMs over Heterogeneous Cyber Threat Intelligence

Model ReleasesDGX agent

arXiv:2510.11974v2 Announce Type: replace-cross Abstract: Cyber Threat Intelligence (CTI) is foundational to modern cybersecurity, enabling organizations to proactively defend against evolving threats

CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe

HardwareDGX agent

arXiv:2604.01489v2 Announce Type: replace-cross Abstract: High-performance GPU kernels are critical to modern machine learning systems, yet developing them remains a manual, expert-driven process. Rec

DAST: A VLM-LLM Framework for Cross-Interface Anomaly Detection in O-RAN

AgentsDGX agent

arXiv:2606.06261v1 Announce Type: cross Abstract: O-RAN enables a disaggregated baseband stack with programmable functions that communicate over standardized open interfaces. The same openness that en

Data Flow Control: Data Safety Policies for AI Agents

Model ReleasesDGX agent

arXiv:2606.05679v1 Announce Type: cross Abstract: Agents increasingly generate SQL, orchestrate pipelines, and automate data analysis on behalf of users. While recent work improves query correctness,

Deciphering Two Training Clocks in Grokking via Deep Linear Network Theory with Conditional ReLU Reduction

ResearchDGX agent

arXiv:2606.05863v1 Announce Type: cross Abstract: Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales. We formalize this phenomeno

Design a Reliable LLM-Integrated Interface for Mortality Forecasting

Local AiDGX agent

arXiv:2606.06235v1 Announce Type: cross Abstract: Mortality forecasting plays an important role in actuarial and policy decision-making, but its implementation remains technically complex and inaccess

Detecting Perspective Shifts in Multi-agent Systems

AgentsDGX agent

arXiv:2512.05013v2 Announce Type: replace Abstract: Generative models augmented with external tools and update mechanisms (or extit{agents}) have demonstrated capabilities beyond intelligent prompting

Differentiable Efficient Operator Search

SafetyDGX agent

arXiv:2606.05232v1 Announce Type: cross Abstract: Efficient multimodal foundation models often rely on manually designed token-reduction operators, such as pruning, merging, pooling, and adaptive rewe

Dimensionality Reduction for Cyberattack Classification: A Comparative Evaluation of PCA and Linear Predictive Coding

ResearchDGX agent

arXiv:2606.05584v1 Announce Type: cross Abstract: High-dimensional feature representations are widely used in machine learning-based cyberattack detection systems. However, they increase computational

Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows

Model ReleasesDGX agent

arXiv:2606.05670v1 Announce Type: new Abstract: Does adding more agents help an LLM workflow once compared systems share the same benchmark loader, tool access, answer contract, usage accounting, and

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss

SafetyDGX agent

arXiv:2606.06418v1 Announce Type: cross Abstract: Many modern applications of deep learning involve training a neural network via a one-step prediction loss (e.g., L^2 regression, cross-entropy), but

DPBench: Structural Determinants of Multi-Agent LLM Coordination Under Simultaneous Resource Contention

Model ReleasesDGX agent

arXiv:2602.13255v2 Announce Type: replace Abstract: We present DPBench, a benchmark for evaluating coordination in multi-agent systems built from large language models. Existing benchmarks measure tas

DragOn: A Benchmark and Dataset for Drag-Based GUI Interactions

Model ReleasesDGX agent

arXiv:2606.06322v1 Announce Type: new Abstract: GUI agents - vision-based models that control desktops, web browsers, and mobile devices through graphical user interfaces - promise to automate a wide

ECI: Effective Contrastive Information to Evaluate Hard-Negatives

Local AiDGX agent

arXiv:2603.20990v2 Announce Type: replace-cross Abstract: Hard-negative source selection for dense retrieval is usually decided only after fine-tuning and downstream evaluation. We propose Effective C

Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing

Model ReleasesDGX agent

arXiv:2606.05950v1 Announce Type: new Abstract: Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models. However, most existing methods remain con

EEGDancer: Dynamic Emotion Latent Space Masked Modeling with Reinforcement Learning for EEG Continuous Emotion Prediction

TutorialsDGX agent

arXiv:2606.05855v1 Announce Type: cross Abstract: Continuous electroencephalography (EEG) emotion prediction aims to model the temporal evolution of human emotional states from EEG signals. Unlike con

Efficient Asynchronous Federated Evaluation with Strategy Similarity Awareness for Intent-Based Networking in Industrial Internet of Things

Local AiDGX agent

arXiv:2512.20627v2 Announce Type: replace-cross Abstract: Intent-Based Networking (IBN) offers a promising paradigm for intelligent and automated network control in Industrial Internet of Things (IIoT

Enhancing Software Engineering Through Closed-Loop Memory Optimization

Model ReleasesDGX agent

arXiv:2606.05646v1 Announce Type: cross Abstract: Large language models (LLMs) have enabled powerful software engineering (SE) agents capable of navigating complex codebases and resolving real-world i

Escaping the Verifier: Learning to Reason via Demonstrations

SafetyDGX agent

arXiv:2511.21667v4 Announce Type: replace-cross Abstract: Training Large Language Models (LLMs) to reason often relies on Reinforcement Learning (RL) with task-specific verifiers. However, many real-w

Evaluating Agentic Configuration Repair for Computer Networks

Model ReleasesDGX agent

arXiv:2606.06212v1 Announce Type: new Abstract: Misconfigurations in computer networks remain a major source of critical Internet outages. Research is turning to Large Language Models (LLMs) to automa

Evaluation of LLMs for Mathematical Formalization in Lean

Model ReleasesDGX agent

arXiv:2606.05632v1 Announce Type: new Abstract: Within the past few years, the ability of Large Language Models (LLMs) to generate formal mathematical proofs has improved drastically. We provide a com

Explainable AI-Driven Cyber Risk Analytics and Model Reliability Assessment for Intelligent Governance of U.S. Critical Infrastructure: An XGBoost and SHAP-Based Intrusion Detection Framework

ApplicationsDGX agent

arXiv:2606.05710v1 Announce Type: cross Abstract: The increasing penetrations of the critical infrastructure sector in the United States with intelligent digital technologies have greatly increased ex

Exploring LLMs for South Asian Music Understanding and Generation

Model ReleasesDGX agent

arXiv:2606.05522v1 Announce Type: cross Abstract: Recent advancements in Large Language Models (LLMs) have shown promising results in music understanding and generation tasks. However, existing works

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

ResearchDGX agent

arXiv:2606.06357v1 Announce Type: cross Abstract: Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio

FIDES: Faithful Inference via Deep Evidence Signals for Retrieval-Memory Conflict in RAG

SafetyDGX agent

arXiv:2606.05644v1 Announce Type: new Abstract: When retrieved evidence contradicts parametric memory, language models frequently ignore context and default to memorized priors -- a failure that under

Finite Element-Based Material Learning via Automatic Differentiation: Learning constitutive neural network models from full-field deformation data

ResearchDGX agent

arXiv:2606.05199v1 Announce Type: cross Abstract: The identification of constitutive neural network models from heterogeneous full-field deformation data provides a robust alternative to traditional c

Fix the Mind, Not the Move: Interpretable AI Assistance via Knowledge-Gap Localization

ResearchDGX agent

arXiv:2606.05602v1 Announce Type: new Abstract: AI assistants in human-AI collaboration often correct suboptimal human actions through behavioral feedback (e.g., alerts or steering-wheel nudges in ass

From Attack Simulation to SIEM Rule: Deterministic Detection-as-Code Synthesis with Probe-Level Traceability

ApplicationsDGX agent

arXiv:2606.05252v1 Announce Type: cross Abstract: Security teams routinely simulate attacks against their own systems to check whether their monitoring would catch a real intruder. These Breach-and-At

From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents

SafetyDGX agent

arXiv:2606.06223v1 Announce Type: new Abstract: Language-model agents act through repeated cycles of observation, reasoning, and action selection, making safety monitoring depend on both internal mode

From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents

SafetyDGX agent

arXiv:2606.05805v1 Announce Type: new Abstract: LLM-based guardrails typically safeguard agents by evaluating proposed actions or inputs before execution, producing safety signals such as binary allow

GenTI: Benchmarking LLMs for Autonomous IDPS Rule Generation for Unseen Attacks

Model ReleasesDGX agent

arXiv:2606.05844v1 Announce Type: cross Abstract: Rule-based Intrusion Detection and Prevention Systems (IDPS) offer precise attack detection as well as mitigation, however their manually crafted, sig

Geographic Bias and Diversity in AI Evaluation

Model ReleasesDGX agent

arXiv:2606.05187v1 Announce Type: cross Abstract: Among the many challenges hindering the responsible development and deployment of AI, arguably none has faced more intense scrutiny than bias in its v

GIPO: Gaussian Importance Sampling Policy Optimization

SafetyDGX agent

arXiv:2603.03955v2 Announce Type: replace-cross Abstract: Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitation.

GITCO: Gated Inference-Time Context Optimization in TSFMs

Model ReleasesDGX agent

arXiv:2606.05332v1 Announce Type: new Abstract: Patch-based Time Series Foundation Models (TSFMs) suffer from context poisoning: structurally anomalous patches capture disproportionate attention and s

Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement

Model ReleasesDGX agent

arXiv:2606.06468v1 Announce Type: new Abstract: We introduce Goedel-Architect, an agentic framework for formal theorem proving in Lean 4 centered on blueprint generation and refinement. A blueprint is

GOTabPFN: From Feature Ordering to Compact Tokenization for Tabular Foundation Models on High-Dimensional Data

TutorialsDGX agent

arXiv:2606.05441v1 Announce Type: cross Abstract: We investigate how to make small tabular foundation models effective for High-Dimensional, Low-Sample Size (HDLSS) tabular prediction without retraini

Gradient descent at the Edge of Stability: free energy model and kinetic description of the two-layer network

ResearchDGX agent

arXiv:2606.05326v1 Announce Type: cross Abstract: We study the dynamics of gradient descent in the Edge of Stability regime, where the learning rate is large enough to induce persistent oscillations i

Gradient Descent with Large Step Size Restores Symmetry in Deep Linear Networks with Multi-Pathway

ResearchDGX agent

arXiv:2606.05219v1 Announce Type: cross Abstract: Recent analyses of multi-pathway Deep Linear Networks use Gradient Flow to predict a 'winner-takes-all' specialization in which path symmetry breaks a

GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection

Model ReleasesDGX agent

arXiv:2606.05566v1 Announce Type: new Abstract: Large Language Models (LLMs) have transformed natural language processing, but they remain vulnerable to Prompt Injection (PI) and Jailbreak (JB) attack

How Far Did They Go? The Persuasive Tactics of Covert LLM Agents in a Discontinued Field Experiment

Model ReleasesDGX agent

arXiv:2606.05256v1 Announce Type: new Abstract: This study analyzes a publicly released dataset from a discontinued field experiment on Reddit's r/ChangeMyView. The intervention, conducted by unknown,

Human Oversight and Overload: Two Hidden and Costly Burdens of AI-Assisted Software Engineering

ResearchDGX agent

arXiv:2606.05770v1 Announce Type: cross Abstract: AI is changing how software engineers work, but it often comes with hidden burdens and costs. In this paper, we characterize two such often-overlooked

Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents

AgentsDGX agent

arXiv:2606.05391v1 Announce Type: cross Abstract: Autonomous software agents hold promise to increase developer productivity but make mistakes and exhibit novel failure modes, making human oversight c

HypRAG: Hyperbolic Dense Retrieval for Retrieval Augmented Generation

SafetyDGX agent

arXiv:2602.07739v2 Announce Type: replace-cross Abstract: Embedding geometry plays a fundamental role in retrieval quality, yet dense retrievers for retrieval-augmented generation (RAG) remain largely

I Know What You Meme, Even If it Emerged Today: Understanding Evolving Memes through Open-World Knowledge Acquisition

Model ReleasesDGX agent

arXiv:2606.05316v1 Announce Type: new Abstract: Multimodal memes are dynamic and often require up to date background knowledge for interpretation. Existing methods often overlook such knowledge or rel

In-Training Defenses against Emergent Misalignment in Language Models

TutorialsDGX agent

arXiv:2508.06249v3 Announce Type: replace-cross Abstract: Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (

Individual Gain, Collective Loss: Metacognitive Adaptation in AI-Assisted Creativity

ResearchDGX agent

arXiv:2606.05532v1 Announce Type: new Abstract: Recent studies reveal a paradox: AI enhances individual creative outputs while reducing collective diversity. Current explanations -- cognitive offloadi

← Previous
1…155156157158159…358
Next →