AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,474 results
9 Jun 2026

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.08094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation-class GPU, a mismatch for the hardware

Voting Protocols as Coordination Mechanisms for Role-Constrained Multi-Agent Tutoring Systems

AgentsDGX agent

arXiv:2606.08030v1 Announce Type: cross Abstract: Agentic tutoring systems introduce a coordination challenge: multiple agents may propose different but reasonable interventions, yet only one response

Weak-Driven Learning: How Weak Agents make Strong Agents Stronger

TutorialsDGX agent

arXiv:2602.08222v2 Announce Type: replace Abstract: As post-training optimization becomes central to improving large language models, we observe a persistent saturation bottleneck: once models grow hi


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

Model ReleasesDGX agent

arXiv:2606.09426v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and ext

Web Agents Should Use Typed Actions Instead of Click-Based Browsing

SafetyDGX agent

arXiv:2602.17245v2 Announce Type: replace Abstract: This position paper argues that building a reliable agentic Web requires shifting from low-level interaction primitives to typed actions supported b

What Makes a Desired Graph for Relational Deep Learning?

SafetyDGX agent

arXiv:2606.08491v1 Announce Type: new Abstract: Relational deep learning (RDL) converts relational databases (RDBs) into heterogeneous graphs, but graphs derived directly from database schemas are oft

What Makes Video World Model Latents Action-Relevant: Prediction over Reconstruction

ResearchDGX agent

arXiv:2606.07687v1 Announce Type: cross Abstract: Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-

What's the Point? Spatial Grammar & Index Resolution for Sign Language Processing

ResearchDGX agent

arXiv:2606.08056v1 Announce Type: cross Abstract: Sign language models are predominantly trained with gloss-sequence or text supervision, thereby under-modeling non-lexical and productive construction

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Model ReleasesDGX agent

arXiv:2606.08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these eva

When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents

Model ReleasesDGX agent

arXiv:2602.08235v2 Announce Type: replace-cross Abstract: Although computer-use agents (CUAs) hold significant potential to automate increasingly complex OS workflows, they can demonstrate unsafe unin

When Does Delegation Beat Majority? A Delegation-Based Aggregator for Multi-Sample LLM Inference

ResearchDGX agent

arXiv:2606.08098v1 Announce Type: new Abstract: Majority voting over sampled answers is the dominant unsupervised aggregator for multi-sample LLM inference. We show that piping the signals every sampl

When No Answer Is Correct: Diagnosing Absent Answer Detection for MLLMs in Video Understanding

ResearchDGX agent

arXiv:2606.08239v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made substantial advancements in video understanding, yet the reliability of their responses remains under

When Video Misreads: Closed-Loop Distillation of Reading Heuristics for Exploratory Manipulation Trace QA

AgentsDGX agent

arXiv:2606.08542v1 Announce Type: cross Abstract: Exploratory manipulation often turns an apparent failed attempt into the key evidence for what to do next. For example, a robot pulls a locked cabinet

Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models

Model ReleasesDGX agent

arXiv:2606.07808v1 Announce Type: new Abstract: Reasoning language models deployed in agentic workflows must follow an instruction hierarchy: when instructions from different sources conflict, the mod

WhiFlash: Accelerating Speculative Decoding with Token-Level Cross-Paradigm Routing

SafetyDGX agent

arXiv:2606.07710v1 Announce Type: cross Abstract: The autoregressive nature of large language models (LLMs) remains a significant bottleneck for inference, particularly in complex agentic workloads. W

Who Earns the Safety? Intervention-Aware Quantum Predictive Control with Safety Attribution

Model ReleasesDGX agent

arXiv:2606.09778v1 Announce Type: cross Abstract: Hard safety filters are increasingly placed downstream of learned controllers to guarantee constraint satisfaction at run time. Yet a filtered control

Why Limit the Residual Stream to Layers and Not Tokens? Persistent Memory for Continuous Latent Reasoning

ResearchDGX agent

arXiv:2606.07720v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable reasoning abilities on mathematical and multi-hop planning tasks. The CoCoNuT (Chain of Contin

XAInomaly: Explainable and Interpretable Deep Contractive Autoencoder for O-RAN Traffic Anomaly Detection

TutorialsDGX agent

arXiv:2502.09194v1 Announce Type: cross Abstract: Generative Artificial Intelligence (AI) techniques have become integral part in advancing next generation wireless communication systems by enabling s

XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad

Model ReleasesDGX agent

arXiv:2601.14063v2 Announce Type: replace-cross Abstract: Cross-cultural competence in large language models (LLMs) requires understanding and adapting Culture-Specific Items (CSIs) across varying cul

Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline

Model ReleasesDGX agent

arXiv:2606.07965v1 Announce Type: new Abstract: Large Visual Language Models (LVLMs) have achieved remarkable success in vision tasks. However, the significant differences between industrial and natur

ZIPP:Zero-shot Image Personalization from Personas

Model ReleasesDGX agent

arXiv:2606.08841v1 Announce Type: new Abstract: Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate a

8 Jun 2026

A Comprehensive Anatomy of Human and DeepSeek-R1 LLM Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2606.07410v1 Announce Type: cross Abstract: The emergence of 'Aha moments' in large language models, particularly DeepSeek-R1-0120, has raised the question of whether these systems genuinely rea

A Geometric Account of Activation Steering through Angle-Norm Decomposition

SafetyDGX agent

arXiv:2606.06735v1 Announce Type: new Abstract: Linear activation steering has gained popularity as a simple and empirically effective way to control language model behavior. More recently, spherical

A Geometric Gaussian Mixture Representation of Plane Curves

Model ReleasesDGX agent

arXiv:2606.06505v1 Announce Type: cross Abstract: We introduce a user defined probabilistic polygonal representation for plane curves. Given a curve, we select vertices on the curve and connect consec

A Geometric View for Understanding Concept Learning and Neuron Interpretation in Sparse Autoencoders

SafetyDGX agent

arXiv:2606.07007v1 Announce Type: cross Abstract: We propose a unified mathematical framework for a geometric understanding of concept learning and neuron interpretation in sparse autoencoders (SAEs).

A Mechanism-Coupled Split Window Network for Medium- to High-Resolution Land Surface Temperature Retrieval

TutorialsDGX agent

arXiv:2509.04991v2 Announce Type: replace-cross Abstract: Land surface temperature (LST) is a fundamental physical variable in land-atmosphere interactions, surface energy budgets, and climate process

A robust PPG foundation model using multimodal physiological supervision

TutorialsDGX agent

arXiv:2606.07365v1 Announce Type: cross Abstract: Photoplethysmography (PPG), a non-invasive measure of changes in blood volume, is widely used in both wearable devices and clinical settings. Recent P

A Study of Parallel Continuous Local Search

ResearchDGX agent

arXiv:2606.06656v1 Announce Type: new Abstract: We study parallel Continuous Local Search (CLS) as a solution approach for Boolean satisfiability problems with symmetric pseudo-Boolean (PB) constraint

A Temporal Spatial Minimax Rate for Smoothly-Varying Distributions in Wasserstein Space

Local AiDGX agent

arXiv:2606.07325v1 Announce Type: cross Abstract: We study the minimax rate of estimating a future value mu_{t_n+h} of a curve tmapstomu_t in the 2-Wasserstein space P_2(R^d) from finitely many noisy

Accelerated Fourier SAT (AFSAT): Fully Realising a GPU-based Symmetric Pseudo-Boolean SAT Solver

HardwareDGX agent

arXiv:2606.06641v1 Announce Type: new Abstract: We present Accelerated Fourier SAT (AFSAT), a GPU-accelerated solver for pseudo-Boolean satisfiability based on continuous local search (CLS). AFSAT rea

Accounting for Context: Shaping Moral Credences for Value Alignment

SafetyDGX agent

arXiv:2606.06972v1 Announce Type: new Abstract: Ensuring that agent behaviours are aligned with human moral values inevitably raises the problem of how to account for the plurality of moral perspectiv

Acoustic Cue Alignment in Audio Language Models for Speech Emotion Recognition

SafetyDGX agent

arXiv:2606.07309v1 Announce Type: cross Abstract: Instruction-following audio language models (ALMs) can be augmented with explicit acoustic cues, yet it remains unclear whether such cues are used in

Act As a Real Researcher: A Suite of Benchmarks Evaluating Frontier LLMs and Agentic Harnesses in Research Lifecycle

Model ReleasesDGX agent

arXiv:2606.07462v1 Announce Type: new Abstract: As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-

AdMem: Advanced Memory for Task-solving Agents

AgentsDGX agent

arXiv:2606.06787v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise as tool-using agents but remain limited in long-horizon tasks that require remembering, organizing, and reusin

AEGIS: A Backup Reflex for Physical AI

SafetyDGX agent

arXiv:2606.06660v1 Announce Type: new Abstract: Long-horizon robot manipulation tends to fail gradually: one bad step degrades the state, and the policy spirals into a basin from which it cannot recov

Agentic Large Language Models for Automated Structural Analysis of 3D Frame Systems

AgentsDGX agent

arXiv:2606.06525v1 Announce Type: cross Abstract: Large language models (LLMs) have emerged as powerful foundation models with strong reasoning capabilities across domains. Beyond reactive text genera

AI-Driven Test Case Generation from Natural Language Requirements: A Survey of Techniques and Research Gaps

ResearchDGX agent

arXiv:2606.06563v1 Announce Type: cross Abstract: Software testing is critical for verifying that systems meet specified requirements, yet remains among the most time-consuming and expensive activitie

AI Sovereignty: A Qualitative Model of Strategic Competition as AI Becomes an Instrument of National Power

ResearchDGX agent

arXiv:2606.07245v1 Announce Type: cross Abstract: AI sovereignty is the extent to which a nation independently controls its artificial intelligence (AI) technologies. The race toward ever-more-sophist

An Abstract Architecture for Explainable Autonomy in Hazardous Environments

SafetyDGX agent

arXiv:2606.07211v1 Announce Type: cross Abstract: Autonomous robotic systems are being proposed for use in hazardous environments, often to reduce the risks to human workers. In the immediate future,

Analysing Differences in Persuasive Language in LLM-Generated Text: Uncovering Stereotypical Gender Patterns

ResearchDGX agent

arXiv:2601.05751v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for everyday communication tasks, including drafting interpersonal messages intended to inf

Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety

SafetyDGX agent

arXiv:2606.06529v1 Announce Type: new Abstract: An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework f

Attention Consistent Longitudinal Medical Visual Question Answering Guided by Vision Foundation Models

Model ReleasesDGX agent

arXiv:2606.06534v1 Announce Type: cross Abstract: Longitudinal medical visual question answering (VQA) requires reasoning about anatomical differences between an image of a current time point and an i

Attention-Guided Autoencoder Fusion for Insulator Defect Detection Using UAV Transmission-Line Imaging

Local AiDGX agent

arXiv:2606.06536v1 Announce Type: cross Abstract: Automated defect detection in high-voltage transmission-line insulators remains challenging due to severe class imbalance, large scale variation, and

Auditing Training Data in Domain-adapted LLMs: LoRA-MINT

Model ReleasesDGX agent

arXiv:2606.06946v1 Announce Type: cross Abstract: We present LoRA-MINT, a new methodology for Membership Inference Test (MINT) applied to recent Large Language Models (LLMs) fine-tuned for specific Na

Autonomous computational catalysis through an agentic research system

AgentsDGX agent

arXiv:2601.13508v4 Announce Type: replace-cross Abstract: Autonomous agents are beginning to transform scientific research from tool-assisted workflows toward self-sustaining discovery processes. Comp

Autonomous heterogeneous catalyst discovery with a self-evolving multi-agent digital twin

Model ReleasesDGX agent

arXiv:2606.05050v1 Announce Type: cross Abstract: Theoretical heterogeneous catalysis promises rapid catalyst discovery, yet computational and machine-learning predictions often deviate from experimen

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation

ApplicationsDGX agent

arXiv:2606.06761v1 Announce Type: cross Abstract: Visuomotor manipulation policies trained via large-scale behavior cloning have achieved strong semantic scene understanding, yet often fail to reliabl

Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio

Model ReleasesDGX agent

arXiv:2603.08683v2 Announce Type: replace-cross Abstract: Autoregressive 'language' models (LMs) trained on raw waveforms can be repurposed for lossless audio compression, but prior work is limited to

Beyond Post-hoc Explanation: Toward Glassbox AI via Probabilistic Mediation

SafetyDGX agent

arXiv:2606.07113v1 Announce Type: new Abstract: Large language models are rapidly becoming infrastructural components in high-stakes institutional settings, including public administration, legal reas

Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy

SafetyDGX agent

arXiv:2606.06903v1 Announce Type: cross Abstract: Human image animation aims to generate a video from a static reference image, guided by pose information extracted from a driving video. Existing appr

Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2606.07244v1 Announce Type: cross Abstract: Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions while navigating in real-world-

Bounded-Abstention Pairwise Learning to Rank

SafetyDGX agent

arXiv:2505.23437v2 Announce Type: replace-cross Abstract: Ranking systems influence decision-making in high-stakes domains like health, education, and employment, where they can have substantial econo

Breaking the Lock-in: Diversifying Text-to-Image Generation via Representation Modulation

SafetyDGX agent

arXiv:2606.06813v1 Announce Type: cross Abstract: Recent text-to-image models built on large-scale Transformer backbones and flow-based objectives deliver strong text-image alignment and high visual q

CAF-Gen: A Multi-Agent System for Enriching Argumentation Structures

SafetyDGX agent

arXiv:2606.06646v1 Announce Type: cross Abstract: Formalizing complex reasoning from natural text is one of the central challenges in computational linguistics. It requires systems to understand not j

CARVE-Q: Quantum-Proposed, Classically Certified Interactive Driving Repair

SafetyDGX agent

arXiv:2606.06531v1 Announce Type: new Abstract: The critical question after a correct driving veto is not only whether a maneuver is unsafe, but whether the blocked interaction admits a lawful, audita

Chameleon: Control-Indexed Prospective Memory for Visuomotor Manipulation

Model ReleasesDGX agent

arXiv:2603.24576v2 Announce Type: replace-cross Abstract: Robots often observe information that determines a future action long before that action is executed. In a shell game, for example, a robot fi

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces

ResearchDGX agent

arXiv:2606.06840v1 Announce Type: cross Abstract: Modern reasoning models offer surprisingly strong zero-shot performance on challenging multi-label tasks that require selecting a small set of relevan

CHDP: Cooperative Hybrid Diffusion Policies for Reinforcement Learning in Parameterized Action Space

SafetyDGX agent

arXiv:2601.05675v2 Announce Type: replace Abstract: Hybrid action space, which combines discrete choices and continuous parameters, is prevalent in domains such as robot control and game AI. However,

ChemQuests: A Curated Chemistry Question-Answer Database Extracted from ChemRxiv papers

ApplicationsDGX agent

arXiv:2505.05232v3 Announce Type: replace Abstract: The rapid expansion of chemistry literature poses significant challenges for researchers seeking to efficiently access domain-specific knowledge. To

ChronoForest: Closed-Loop Multi-Tree Diffusion Planning for Efficient Bridge Search and Route Composition

ResearchDGX agent

arXiv:2606.06618v1 Announce Type: cross Abstract: How can we plan long-horizon routes that reach designated goals, visit required waypoints, and remain short when only short-horizon offline trajectori

← Previous
1…150151152153154…358
Next →