AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
20,981 results
14 Apr 2026

THEIA: Learning Complete Kleene Three-Valued Logic in a Pure-Neural Modular Architecture

Model ReleasesDGX agent

arXiv:2604.11284v1 Announce Type: cross Abstract: We present THEIA, a modular neural architecture that learns complete Kleene three-valued logic (K3) end-to-end without any external symbolic solver, a

Think Before you Write: QA-Guided Reasoning for Character Descriptions in Books

ResearchDGX agent

arXiv:2604.11435v1 Announce Type: cross Abstract: Character description generation is an important capability for narrative-focused applications such as summarization, story analysis, and character-dr

Think in Sentences: Explicit Sentence Boundaries Enhance Language Model's Capabilities

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.10135v1 Announce Type: cross Abstract: Researchers have explored different ways to improve large language models (LLMs)' capabilities via dummy token insertion in contexts. However, existin

Think Parallax: Solving Multi-Hop Problems via Multi-View Knowledge-Graph-Based Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2510.15552v3 Announce Type: replace-cross Abstract: Large language models (LLMs) still struggle with multi-hop reasoning over knowledge-graphs (KGs), and we identify a previously overlooked stru

Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation

Model ReleasesDGX agent

arXiv:2604.10511v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for causal and counterfactual reasoning, yet their reliability in real-world policy evaluation remain

Thought Branches: Interpreting LLM Reasoning Requires Resampling

SafetyDGX agent

arXiv:2510.27484v2 Announce Type: replace-cross Abstract: Most work interpreting reasoning models studies only a single chain-of-thought (CoT), yet these models define distributions over many possible

Three Roles, One Model: Role Orchestration at Inference Time to Close the Performance Gap Between Small and Large Agents

Model ReleasesDGX agent

arXiv:2604.11465v1 Announce Type: new Abstract: Large language model (LLM) agents show promise on realistic tool-use tasks, but deploying capable agents on modest hardware remains challenging. We stud

Time is Not a Label: Continuous Phase Rotation for Temporal Knowledge Graphs and Agentic Memory

Model ReleasesDGX agent

arXiv:2604.11544v1 Announce Type: cross Abstract: Structured memory representations such as knowledge graphs are central to autonomous agents and other long-lived systems. However, most existing appro

TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance

TutorialsDGX agent

arXiv:2509.26627v2 Announce Type: replace Abstract: Designing dense rewards is crucial for reinforcement learning (RL), yet in robotics it often demands extensive manual effort and lacks scalability.

TimeSeriesExamAgent: Creating Time Series Reasoning Benchmarks at Scale

Model ReleasesDGX agent

arXiv:2604.10291v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown promising performance in time series modeling tasks, but do they truly understand time series data? While multip

TInR: Exploring Tool-Internalized Reasoning in Large Language Models

SafetyDGX agent

arXiv:2604.10788v1 Announce Type: cross Abstract: Tool-Integrated Reasoning (TIR) has emerged as a promising direction by extending Large Language Models' (LLMs) capabilities with external tools durin

Tipiano: Cascaded Piano Hand Motion Synthesis via Fingertip Priors

TutorialsDGX agent

arXiv:2604.09692v1 Announce Type: new Abstract: Synthesizing realistic piano hand motions requires both precision and naturalness. Physics-based methods achieve precision but produce stiff motions; da

Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference

Model ReleasesDGX agent

arXiv:2604.09613v1 Announce Type: cross Abstract: Production vLLM fleets provision every instance for worst-case context length, wasting 4-8x concurrency on the 80-95% of requests that are short and s

TokUR: Token-Level Uncertainty Estimation for Large Language Model Reasoning

ResearchDGX agent

arXiv:2505.11737v4 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have demonstrated impressive capabilities, their output quality remains inconsistent across various applica

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models

Model ReleasesDGX agent

arXiv:2604.10733v1 Announce Type: cross Abstract: Large language models increasingly serve as conversational agents that adopt personas and role-play characters at user request. This capability, while

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training

Model ReleasesDGX agent

arXiv:2604.10784v1 Announce Type: new Abstract: Recent advances in unified multimodal models (UMMs) have led to a proliferation of architectures capable of understanding, generating, and editing acros

Toward Accountable AI-Generated Content on Social Platforms: Steganographic Attribution and Multimodal Harm Detection

ResearchDGX agent

arXiv:2604.10460v1 Announce Type: cross Abstract: The rapid growth of generative AI has introduced new challenges in content moderation and digital forensics. In particular, benign AI-generated images

Towards Adaptive Open-Set Object Detection via Category-Level Collaboration Knowledge Mining

SafetyDGX agent

arXiv:2604.11195v1 Announce Type: cross Abstract: Existing object detectors often struggle to generalize across domains while adapting to emerging novel categories. Adaptive open-set object detection

Towards an Appropriate Level of Reliance on AI: A Preliminary Reliance-Control Framework for AI in Software Engineering

ResearchDGX agent

arXiv:2604.10530v1 Announce Type: cross Abstract: How software developers interact with Artificial Intelligence (AI)-powered tools, including Large Language Models (LLMs), plays a vital role in how th

Towards Automated Solar Panel Integrity: Hybrid Deep Feature Extraction for Advanced Surface Defect Identification

ResearchDGX agent

arXiv:2604.10969v1 Announce Type: cross Abstract: To ensure energy efficiency and reliable operations, it is essential to monitor solar panels in generation plants to detect defects. It is quite labor

Towards Autonomous Mechanistic Reasoning in Virtual Cells

AgentsDGX agent

arXiv:2604.11661v1 Announce Type: cross Abstract: Large language models (LLMs) have recently gained significant attention as a promising approach to accelerate scientific discovery. However, their app

Towards Green Wearable Computing: A Physics-Aware Spiking Neural Network for Energy-Efficient IMU-based Human Activity Recognition

ResearchDGX agent

arXiv:2604.10458v1 Announce Type: cross Abstract: Wearable IMU-based Human Activity Recognition (HAR) relies heavily on Deep Neural Networks (DNNs), which are burdened by immense computational and buf

Towards Proactive Information Probing: Customer Service Chatbots Harvesting Value from Conversation

ResearchDGX agent

arXiv:2604.11077v1 Announce Type: new Abstract: Customer service chatbots are increasingly expected to serve not merely as reactive support tools for users, but as strategic interfaces for harvesting

Towards Reasonable Concept Bottleneck Models

ResearchDGX agent

arXiv:2506.05014v2 Announce Type: replace-cross Abstract: We propose a novel, flexible, and efficient framework for designing Concept Bottleneck Models (CBMs) that enables practitioners to explicitly

Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs

Model ReleasesDGX agent

arXiv:2604.10480v1 Announce Type: new Abstract: Post-training data plays a pivotal role in shaping the capabilities of Large Language Models (LLMs), yet datasets are often treated as isolated artifact

Training Deep Visual Networks Beyond Loss and Accuracy Through a Dynamical Systems Approach

ResearchDGX agent

arXiv:2604.09716v1 Announce Type: cross Abstract: Deep visual recognition models are usually trained and evaluated using metrics such as loss and accuracy. While these measures show whether a model is

TrajOnco: a multi-agent framework for temporal reasoning over longitudinal EHR for multi-cancer early detection

Model ReleasesDGX agent

arXiv:2604.10386v1 Announce Type: new Abstract: Accurate estimation of cancer risk from longitudinal electronic health records (EHRs) could support earlier detection and improved care, but modeling su

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards

Model ReleasesDGX agent

arXiv:2604.10110v1 Announce Type: new Abstract: Large Language Models (LLMs) have become a key foundation for enabling personalized smart home experiences. While existing studies have explored how sma

Tuning Language Models for Robust Prediction of Diverse User Behaviors

ApplicationsDGX agent

arXiv:2505.17682v2 Announce Type: replace-cross Abstract: Predicting user behavior is essential for intelligent assistant services, yet deep learning models often struggle to capture long-tailed behav

Tuning Qwen2.5-VL to Improve Its Web Interaction Skills

Model ReleasesDGX agent

arXiv:2604.09571v1 Announce Type: cross Abstract: Recent advances in vision-language models (VLMs) have sparked growing interest in using them to automate web tasks, yet their feasibility as independe

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization

Model ReleasesDGX agent

arXiv:2604.09574v1 Announce Type: new Abstract: The rise of autonomous GUI agents has triggered adversarial countermeasures from digital platforms, yet existing research prioritizes utility and robust

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization

Model ReleasesDGX agent

arXiv:2604.10721v1 Announce Type: cross Abstract: Natural-language Guided Cross-view Geo-localization (NGCG) aims to retrieve geo-tagged satellite imagery using textual descriptions of ground scenes.

UBio-MolFM: A Universal Molecular Foundation Model for Bio-Systems

Local AiDGX agent

arXiv:2602.17709v2 Announce Type: replace-cross Abstract: All-atom molecular simulation serves as a quintessential ``computational microscope'' for understanding the machinery of life, yet it remains

UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation

ResearchDGX agent

arXiv:2604.10485v1 Announce Type: cross Abstract: Low-visibility scenarios, such as low-light conditions, pose significant challenges to human pose estimation due to the scarcity of annotated low-ligh

Uncertainty-Aware Web-Conditioned Scientific Fact-Checking

ResearchDGX agent

arXiv:2604.11036v1 Announce Type: cross Abstract: Scientific fact-checking is vital for assessing claims in specialized domains such as biomedicine and materials science, yet existing systems often ha

Understanding Generalization in Role-Playing Models via Information Theory

ApplicationsDGX agent

arXiv:2512.17270v2 Announce Type: replace-cross Abstract: Role-playing models (RPMs) are widely used in real-world applications but underperform when deployed in the wild. This degradation can be attr

Unifying Ontology Construction and Semantic Alignment for Deterministic Enterprise Reasoning at Scale

Model ReleasesDGX agent

arXiv:2604.09608v1 Announce Type: new Abstract: While enterprises amass vast quantities of data, much of it remains chaotic and effectively dormant, preventing decision-making based on comprehensive i

Unilateral Relationship Revision Power in Human-AI Companion Interaction

AgentsDGX agent

arXiv:2603.23315v3 Announce Type: replace-cross Abstract: When providers update AI companions, users report grief, betrayal, and loss. A growing literature asks whether the norms governing personal re

UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents

Model ReleasesDGX agent

arXiv:2604.11557v1 Announce Type: new Abstract: Tool-use capability is a fundamental component of LLM agents, enabling them to interact with external systems through structured function calls. However

Universal statistical signatures of evolution in artificial intelligence architectures

ResearchDGX agent

arXiv:2604.10571v1 Announce Type: cross Abstract: We test whether artificial intelligence architectural evolution obeys the same statistical laws as biological evolution. Compiling 935 ablation experi

Unsupervised Detection of Spatiotemporal Anomalies in PMU Data Using Transformer-Based BiGAN

Model ReleasesDGX agent

arXiv:2509.25612v2 Announce Type: replace-cross Abstract: Ensuring power grid resilience requires the timely and unsupervised detection of anomalies in synchrophasor data streams. We introduce T-BiGAN

Use of AI Tools: Guidelines to Maintain Academic Integrity in Computing Colleges

TutorialsDGX agent

arXiv:2604.11111v1 Announce Type: cross Abstract: The rapid adoption of AI tools such as ChatGPT has significantly transformed academic practices, offering considerable benefits for both students and

Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control

Model ReleasesDGX agent

arXiv:2604.03147v2 Announce Type: replace-cross Abstract: We present a method to identify a valence-arousal (VA) subspace within large language model representations. From 211k emotion-labeled texts,

Variance-Aware Prior-Based Tree Policies for Monte Carlo Tree Search

SafetyDGX agent

arXiv:2512.21648v2 Announce Type: replace-cross Abstract: Monte Carlo Tree Search (MCTS) has profoundly influenced reinforcement learning (RL) by integrating planning and learning in tasks requiring l

Variational Visual Question Answering for Uncertainty-Aware Selective Prediction

ResearchDGX agent

arXiv:2505.09591v3 Announce Type: replace-cross Abstract: Despite remarkable progress in recent years, Vision Language Models (VLMs) remain prone to overconfidence and hallucinations on tasks such as

Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis

AgentsDGX agent

arXiv:2604.10800v1 Announce Type: cross Abstract: Learned classifiers deployed in agentic pipelines face a fundamental reliability problem: predictions are probabilistic inferences, not verified concl

VeriSim: A Configurable Framework for Evaluating Medical AI Under Realistic Patient Noise

ResearchDGX agent

arXiv:2604.10441v1 Announce Type: new Abstract: Medical large language models (LLMs) achieve impressive performance on standardized benchmarks, yet these evaluations fail to capture the complexity of

VeriTrans: Fine-Tuned LLM-Assisted NL-to-PL Translation via a Deterministic Neuro-Symbolic Pipeline

SafetyDGX agent

arXiv:2604.10341v1 Announce Type: new Abstract: extbf{VeriTrans} is a reliability-first ML system that compiles natural-language requirements into solver-ready logic with validator-gated reliability.

VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation

ResearchDGX agent

arXiv:2604.02467v2 Announce Type: replace-cross Abstract: Cinematic camera control relies on a tight feedback loop between director and cinematographer, where camera motion and framing are continuousl

VGA-Bench: A Unified Benchmark and Multi-Model Framework for Video Aesthetics and Generation Quality Evaluation

Model ReleasesDGX agent

arXiv:2604.10127v1 Announce Type: cross Abstract: The rapid advancement of AIGC-based video generation has underscored the critical need for comprehensive evaluation frameworks that go beyond traditio

Vibe-driven model-based engineering

ResearchDGX agent

arXiv:2604.10645v1 Announce Type: cross Abstract: There is a pressing need for better development methods and tools to keep up with the growing demand and increasing complexity of new software systems

VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories

Model ReleasesDGX agent

arXiv:2604.10542v1 Announce Type: cross Abstract: Video-to-Audio (V2A) generation is essential for immersive multimedia experiences, yet its evaluation remains underexplored. Existing benchmarks typic

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG

SafetyDGX agent

arXiv:2604.05418v2 Announce Type: replace-cross Abstract: Scaling multimodal large language models (MLLMs) to long videos is constrained by limited context windows. While retrieval-augmented generatio

Virtual Smart Metering in District Heating Networks via Heterogeneous Spatial-Temporal Graph Neural Networks

Model ReleasesDGX agent

arXiv:2604.10166v1 Announce Type: cross Abstract: Intelligent operation of thermal energy networks aims to improve energy efficiency, reliability, and operational flexibility through data-driven contr

Volumetric Ergodic Control

ResearchDGX agent

arXiv:2511.11533v3 Announce Type: replace-cross Abstract: Ergodic control synthesizes optimal coverage behaviors over spatial distributions for nonlinear systems. However, existing formulations model

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments

Model ReleasesDGX agent

arXiv:2506.02387v3 Announce Type: replace Abstract: Recent advancements in Vision Language Models (VLMs) have expanded their capabilities to interactive agent tasks, yet existing benchmarks remain lim

WaveMoE: A Wavelet-Enhanced Mixture-of-Experts Foundation Model for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2604.10544v1 Announce Type: cross Abstract: Time series foundation models (TSFMs) have recently achieved remarkable success in universal forecasting by leveraging large-scale pretraining on dive

WearBCI Dataset: Understanding and Benchmarking Real-World Wearable Brain-Computer Interfaces Signals

Model ReleasesDGX agent

arXiv:2604.09649v1 Announce Type: cross Abstract: Brain-computer interfaces (BCIs) have opened new platforms for human-computer interaction, medical diagnostics, and neurorehabilitation. Wearable BCI

WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark

Model ReleasesDGX agent

arXiv:2604.10988v1 Announce Type: new Abstract: Existing browser agent benchmarks face a fundamental trilemma: real-website benchmarks lack reproducibility due to content drift, controlled environment

WebLLM: A High-Performance In-Browser LLM Inference Engine

Local AiDGX agent

arXiv:2412.15803v2 Announce Type: replace-cross Abstract: Advancements in large language models (LLMs) have unlocked remarkable capabilities. While deploying these models typically requires server-gra

← Previous
1…337338339340341…350
Next →