AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
22 Apr 2026

Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2604.18663v1 Announce Type: cross Abstract: Existing jamming attacks on Retrieval-Augmented Generation (RAG) systems typically induce explicit refusals or denial-of-service behaviors, which are

Beyond Itinerary Planning-A Real-World Benchmark for Multi-Turn and Tool-Using Travel Tasks

Model ReleasesDGX agent

arXiv:2512.22673v3 Announce Type: replace Abstract: Travel planning is a natural real-world task to test large language models' (LLMs) planning and tool-use abilities. Although prior work has studied

Beyond One Output: Visualizing and Comparing Distributions of Language Model Generations

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.18724v1 Announce Type: new Abstract: Users typically interact with and evaluate language models via single outputs, but each output is just one sample from a broad distribution of possible

Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications

SafetyDGX agent

arXiv:2604.19281v1 Announce Type: cross Abstract: The use of Large Language Models (LLMs) to support patients in addressing medical questions is becoming increasingly prevalent. However, most of the m

Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration

SafetyDGX agent

arXiv:2604.17457v2 Announce Type: replace-cross Abstract: Dynamic programming is one of the most fundamental methodologies for solving Markov decision problems. Among its many variants, Q-value iterat

Beyond the 'Diff': Addressing Agentic Entropy in Agentic Software Development

AgentsDGX agent

arXiv:2604.16323v2 Announce Type: replace-cross Abstract: As autonomous coding agents become deeply embedded in software development workflows, their high operational velocity introduces a critical ov

Bootstrapping Code Translation with Weighted Multilanguage Exploration

ResearchDGX agent

arXiv:2601.03512v2 Announce Type: replace-cross Abstract: Code translation across multiple programming languages is essential yet challenging due to two vital obstacles: scarcity of parallel data pair

Bridging the High-Frequency Data Gap: A Millisecond-Resolution Network Dataset for Advancing Time Series Foundation Models

ApplicationsDGX agent

arXiv:2603.16497v2 Announce Type: replace-cross Abstract: Time series foundation models (TSFMs) require diverse, real-world datasets to adapt across varying domains and temporal frequencies. However,

CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark

Model ReleasesDGX agent

arXiv:2505.16968v4 Announce Type: replace-cross Abstract: Cross-architecture GPU code transpilation is essential for unlocking low-level hardware portability, yet no scalable solution exists. We intro

CentaurTA Studio: A Self-Improving Human-Agent Collaboration System for Thematic Analysis

SafetyDGX agent

arXiv:2604.18589v1 Announce Type: cross Abstract: Thematic analysis is difficult to scale: manual workflows are labor-intensive, while fully automated pipelines often lack controllability and transpar

Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models

SafetyDGX agent

arXiv:2511.06168v3 Announce Type: replace Abstract: This paper primarily demonstrates a method to quantitatively assess the alignment between multi-step, structured reasoning in large language models

Characterizing AlphaEarth Embedding Geometry for Agentic Environmental Reasoning

Model ReleasesDGX agent

arXiv:2604.18715v1 Announce Type: cross Abstract: Earth observation foundation models encode land surface information into dense embedding vectors, yet the geometric structure of these representations

Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language

Model ReleasesDGX agent

arXiv:2604.19667v1 Announce Type: cross Abstract: At present, executable visual workflows have emerged as a mainstream paradigm in real-world industrial deployments, offering strong reliability and co

Chimera: Neuro-Symbolic Attention Primitives for Trustworthy Dataplane Intelligence

ResearchDGX agent

arXiv:2602.12851v3 Announce Type: replace-cross Abstract: Deploying expressive learning models directly on programmable dataplanes promises line-rate, low-latency traffic analysis but remains hindered

Choose Your Own Adventure: Non-Linear AI-Assisted Programming with EvoGraph

ResearchDGX agent

arXiv:2604.18883v1 Announce Type: cross Abstract: Current AI-assisted programming tools are predominantly linear and chat-based, which deviates from the iterative and branching nature of programming i

ClawNet: Human-Symbiotic Agent Network for Cross-User Autonomous Cooperation

AgentsDGX agent

arXiv:2604.19211v1 Announce Type: new Abstract: Current AI agent frameworks have made remarkable progress in automating individual tasks, yet all existing systems serve a single user. Human productivi

Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models

SafetyDGX agent

arXiv:2510.26782v3 Announce Type: replace-cross Abstract: A world model is an internal model that simulates how the world evolves. Given past observations and actions, it predicts the future physical

Co-Refine: AI-Powered Tool Supporting Qualitative Analysis

ResearchDGX agent

arXiv:2604.19309v1 Announce Type: cross Abstract: Qualitative coding relies on a researcher's application of codes to textual data. As coding proceeds across large datasets, interpretations of codes o

CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation

ResearchDGX agent

arXiv:2604.19648v1 Announce Type: cross Abstract: SAM3 advances open-vocabulary semantic segmentation by introducing a prompt-driven mask generation paradigm. However, in multi-class open-vocabulary s

CoDA: Towards Effective Cross-domain Knowledge Transfer via CoT-guided Domain Adaptation

ApplicationsDGX agent

arXiv:2604.19488v1 Announce Type: new Abstract: Large language models (LLMs) have achieved substantial advances in logical reasoning, yet they continue to lag behind human-level performance. In-contex

COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition

Local AiDGX agent

arXiv:2503.07259v2 Announce Type: replace-cross Abstract: The goal of creating intelligent, human-centered wearable systems for continuous activity understanding faces a fundamental trade-off: Egocent

Compile to Compress: Boosting Formal Theorem Provers by Compiler Outputs

Model ReleasesDGX agent

arXiv:2604.18587v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated significant potential in formal theorem proving, yet state-of-the-art performance often necessitates pr

Conjuring Semantic Similarity

ResearchDGX agent

arXiv:2410.16431v4 Announce Type: replace Abstract: The semantic similarity between sample expressions measures the distance between their latent 'meaning'. These meanings are themselves typically rep

Council Mode: Mitigating Hallucination and Bias in LLMs via Multi-Agent Consensus

Model ReleasesDGX agent

arXiv:2604.02923v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs), particularly those employing Mixture-of-Experts (MoE) architectures, have achieved remarkable capabilities acros

CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering

Model ReleasesDGX agent

arXiv:2603.16091v2 Announce Type: replace-cross Abstract: In factual question answering, many errors are not failures of access but failures of commitment: the system retrieves relevant evidence, yet

Counting Worlds Branching Time Semantics for post-hoc Bias Mitigation in generative AI

SafetyDGX agent

arXiv:2604.19431v1 Announce Type: cross Abstract: Generative AI systems are known to amplify biases present in their training data. While several inference-time mitigation strategies have been propose

Cross-Model Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Across Three Large Language Models

Model ReleasesDGX agent

arXiv:2604.19598v1 Announce Type: cross Abstract: This study compared repeated generation consistency of exercise prescription outputs across three large language models (LLMs), specifically GPT-4.1,

CulturALL: Benchmarking Multilingual and Multicultural Competence of LLMs on Grounded Tasks

Model ReleasesDGX agent

arXiv:2604.19262v1 Announce Type: cross Abstract: Large language models (LLMs) are now deployed worldwide, inspiring a surge of benchmarks that measure their multilingual and multicultural abilities.

Curiosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model Training

Local AiDGX agent

arXiv:2604.18701v1 Announce Type: cross Abstract: Local prediction-error-based curiosity rewards focus on the current transition without considering the world model's cumulative prediction error acros

Curvature-Aware PCA with Geodesic Tangent Space Aggregation for Semi-Supervised Learning

SafetyDGX agent

arXiv:2604.18816v1 Announce Type: cross Abstract: Principal Component Analysis (PCA) is a fundamental tool for representation learning, but its global linear formulation fails to capture the structure

Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps

Model ReleasesDGX agent

arXiv:2604.19533v1 Announce Type: cross Abstract: We introduce the Cyber Defense Benchmark, a benchmark for measuring how well large language model (LLM) agents perform the core SOC analyst task of th

DanceCrafter: Fine-Grained Text-Driven Controllable Dance Generation via Choreographic Syntax

ResearchDGX agent

arXiv:2604.18648v1 Announce Type: cross Abstract: Text-driven controllable dance generation remains under-explored, primarily due to the severe scarcity of high-quality datasets and the inherent diffi

Debug2Fix: Can Interactive Debugging Help Coding Agents Fix More Bugs?

Model ReleasesDGX agent

arXiv:2602.18571v2 Announce Type: replace-cross Abstract: While significant progress has been made in automating various aspects of software development through coding agents, there is still significa

Decompose, Structure, and Repair: A Neuro-Symbolic Framework for Autoformalization via Operator Trees

Model ReleasesDGX agent

arXiv:2604.19000v1 Announce Type: cross Abstract: Statement autoformalization acts as a critical bridge between human mathematics and formal mathematics by translating natural language problems into f

Decomposed Trust: Privacy, Adversarial Robustness, Ethics, and Fairness in Low-Rank LLMs

SafetyDGX agent

arXiv:2511.22099v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have driven major advances across domains, yet their massive size hinders deployment in resource-constrained sett

Design Rules for Extreme-Edge Scientific Computing on AI Engines

ResearchDGX agent

arXiv:2604.19106v1 Announce Type: cross Abstract: Extreme-edge scientific applications use machine learning models to analyze sensor data and make real-time decisions. Their stringent latency and thro

Detecting Data Contamination in Large Language Models

ResearchDGX agent

arXiv:2604.19561v1 Announce Type: new Abstract: Large Language Models (LLMs) utilize large amounts of data for their training, some of which may come from copyrighted sources. Membership Inference Att

Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps

Model ReleasesDGX agent

arXiv:2604.19565v1 Announce Type: cross Abstract: Hallucinations in Speech Large Language Models (SpeechLLMs) pose significant risks, yet existing detection methods typically rely on gold-standard out

Diamond Maps: Efficient Reward Alignment via Stochastic Flow Maps

SafetyDGX agent

arXiv:2602.05993v2 Announce Type: replace-cross Abstract: Flow and diffusion models produce high-quality samples, but adapting them to user preferences or constraints post-training remains costly and

Distillation Traps and Guards: A Calibration Knob for LLM Distillability

Local AiDGX agent

arXiv:2604.18963v1 Announce Type: cross Abstract: Knowledge distillation (KD) transfers capabilities from large language models (LLMs) to smaller students, yet it can fail unpredictably and also under

Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture The Flag Challenges

Model ReleasesDGX agent

arXiv:2604.19354v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly proposed for autonomous cybersecurity tasks, but their capabilities in realistic offensive settings r

Do LLMs Game Formalization? Evaluating Faithfulness in Logical Reasoning

Model ReleasesDGX agent

arXiv:2604.19459v1 Announce Type: new Abstract: Formal verification guarantees proof validity but not formalization faithfulness. For natural-language logical reasoning, where models construct axiom s

DP-FlogTinyLLM: Differentially private federated log anomaly detection using Tiny LLMs

Model ReleasesDGX agent

arXiv:2604.19118v1 Announce Type: cross Abstract: Modern distributed systems generate massive volumes of log data that are critical for detecting anomalies and cyber threats. However, in real world se

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling

SafetyDGX agent

arXiv:2604.19544v1 Announce Type: new Abstract: Multimodal reward models (MRMs) play a crucial role in aligning Multimodal Large Language Models (MLLMs) with human preferences. Training a good MRM req

DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning

Model ReleasesDGX agent

arXiv:2604.18964v1 Announce Type: new Abstract: This paper introduces DW-Bench, a new benchmark that evaluates large language models (LLMs) on graph-topology reasoning over data warehouse schemas, exp

Early Pruning for Public Transport Routing

ResearchDGX agent

arXiv:2603.12592v2 Announce Type: replace-cross Abstract: Routing algorithms for public transport, particularly the widely used RAPTOR and its variants, often face performance bottlenecks during the t

Easy Samples Are All You Need: Self-Evolving LLMs via Data-Efficient Reinforcement Learning

ResearchDGX agent

arXiv:2604.18639v1 Announce Type: cross Abstract: Previous LLMs-based RL studies typically follow either supervised learning with high annotation costs, or unsupervised paradigms using voting or entro

EgoSelf: From Memory to Personalized Egocentric Assistant

ResearchDGX agent

arXiv:2604.19564v1 Announce Type: cross Abstract: Egocentric assistants often rely on first-person view data to capture user behavior and context for personalized services. Since different users exhib

EHRAG: Bridging Semantic Gaps in Lightweight GraphRAG via Hybrid Hypergraph Construction and Retrieval

ResearchDGX agent

arXiv:2604.17458v2 Announce Type: replace Abstract: Graph-based Retrieval-Augmented Generation (GraphRAG) enhances LLMs by structuring corpus into graphs to facilitate multi-hop reasoning. While recen

Enabling Vibration-Based Gesture Recognition on Everyday Furniture via Energy-Efficient FPGA Implementation of 1D Convolutional Networks

ApplicationsDGX agent

arXiv:2510.23156v2 Announce Type: replace-cross Abstract: The growing demand for smart home interfaces has increased interest in non-intrusive sensing methods like vibration-based gesture recognition.

End-to-End Large Portfolio Optimization for Variance Minimization with Neural Networks through Covariance Cleaning

TutorialsDGX agent

arXiv:2507.01918v3 Announce Type: replace-cross Abstract: We develop a rotation-invariant neural network that provides the global minimum-variance portfolio by jointly learning how to lag-transform hi

Enhancing Construction Worker Safety in Extreme Heat: A Machine Learning Approach Utilizing Wearable Technology for Predictive Health Analytics

SafetyDGX agent

arXiv:2604.19559v1 Announce Type: new Abstract: Construction workers are highly vulnerable to heat stress, yet tools that translate real-time physiological data into actionable safety intelligence rem

Environmental Sound Deepfake Detection Using Deep-Learning Framework

Model ReleasesDGX agent

arXiv:2604.19652v1 Announce Type: cross Abstract: In this paper, we propose a deep-learning framework for environmental sound deepfake detection (ESDD) -- the task of identifying whether the sound sce

Epistemic Skills: Reasoning about Knowledge and Oblivion

ResearchDGX agent

arXiv:2504.01733v4 Announce Type: replace Abstract: This paper presents a class of epistemic logics that captures the dynamics of acquiring knowledge and descending into oblivion, while incorporating

Error-free Training for MedMNIST Datasets

ResearchDGX agent

arXiv:2604.18916v1 Announce Type: new Abstract: In this paper, we introduce a new concept called Artificial Special Intelligence by which Machine Learning models for the classification problem can be

Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks

Model ReleasesDGX agent

arXiv:2604.18660v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in education, yet their default helpfulness often conflicts with pedagogical principles. Prior work

Evaluation-driven Scaling for Scientific Discovery

Local AiDGX agent

arXiv:2604.19341v1 Announce Type: cross Abstract: Language models are increasingly used in scientific discovery to generate hypotheses, propose candidate solutions, implement systems, and iteratively

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale

AgentsDGX agent

arXiv:2604.17406v2 Announce Type: replace Abstract: The convergence of large language models and agents is catalyzing a new era of scientific discovery: Agentic Science. While the scientific method is

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training

SafetyDGX agent

arXiv:2604.19485v1 Announce Type: cross Abstract: Reinforcement learning (RL) for LLM post-training faces a fundamental design choice: whether to use a learned critic as a baseline for policy optimiza

Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models

ResearchDGX agent

arXiv:2604.18786v1 Announce Type: cross Abstract: Scientific feasibility assessment asks whether a claim is consistent with established knowledge and whether experimental evidence could support or ref

← Previous
1…316317318319320…354
Next →