AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
58,761 results
Research

U-CECE: A Universal Multi-Resolution Framework for Conceptual Counterfactual Explanations

DGX agent

arXiv:2604.08295v1 Announce Type: cross Abstract: As AI models grow more complex, explainability is essential for building trust, yet concept-based counterfactual methods still face a trade-off betwee

researcharxiv-cs-cv
10 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding

DGX agent

arXiv:2507.22025v4 Announce Type: replace Abstract: The emergence of Multimodal Large Language Models (MLLMs) has driven significant advances in Graphical User Interface (GUI) agent capabilities. Neve

agentsarxiv-cs-ai
10 Apr 2026
Agents

Uncertainty Estimation for Deep Reconstruction in Actuatic Disaster Scenarios with Autonomous Vehicles

DGX agent

arXiv:2604.06387v1 Announce Type: cross Abstract: Accurate reconstruction of environmental scalar fields from sparse onboard observations is essential for autonomous vehicles engaged in aquatic monito

agentsarxiv-cs-ai
10 Apr 2026
Applications

Understanding Structured Financial Data with LLMs: A Case Study on Fraud Detection

DGX agent

arXiv:2512.13040v2 Announce Type: replace-cross Abstract: Detecting fraud in financial transactions typically relies on tabular models that demand heavy feature engineering to handle high-dimensional

applicationsarxiv-cs-cl
10 Apr 2026
Tutorials

Understanding Task Transfer in Vision-Language Models

DGX agent

arXiv:2511.18787v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) perform well on multimodal benchmarks but lag behind humans and specialized models on visual perception tasks like dep

tutorialsarxiv-cs-cv
10 Apr 2026
Research

Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator

DGX agent

arXiv:2604.08121v1 Announce Type: new Abstract: Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher co

researcharxiv-cs-cv
10 Apr 2026
Model Releases

Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs

DGX agent

arXiv:2601.21463v2 Announce Type: replace-cross Abstract: Existing speech editing detection (SED) datasets are predominantly constructed using manual splicing or limited editing operations, resulting

model-releasesarxiv-cs-ai
10 Apr 2026
Applications

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models

DGX agent

arXiv:2602.20231v2 Announce Type: replace-cross Abstract: Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-acti

applicationsarxiv-cs-cv
10 Apr 2026
Model Releases

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding

DGX agent

arXiv:2604.08522v1 Announce Type: new Abstract: Video temporal grounding (VTG) is typically tackled with dataset-specific models that transfer poorly across domains and query styles. Recent efforts to

model-releasesarxiv-cs-cv
10 Apr 2026
Applications

Unsupervised Neural Network for Automated Classification of Surgical Urgency Levels in Medical Transcriptions

DGX agent

arXiv:2604.06214v1 Announce Type: cross Abstract: Efficient classification of surgical procedures by urgency is paramount to optimize patient care and resource allocation within healthcare systems. Th

applicationsarxiv-cs-ai
10 Apr 2026
Safety

URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection

DGX agent

arXiv:2604.06728v1 Announce Type: cross Abstract: Multimodal sarcasm detection (MSD) aims to identify sarcastic intent from semantic incongruity between text and image. Although recent methods have im

safetyarxiv-cs-ai
10 Apr 2026
Model Releases

Validated Intent Compilation for Constrained Routing in LEO Mega-Constellations

DGX agent

arXiv:2604.07264v1 Announce Type: cross Abstract: Operating LEO mega-constellations requires translating high-level operator intents ('reroute financial traffic away from polar links under 80 ms') int

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents

DGX agent

arXiv:2603.15118v2 Announce Type: replace Abstract: We introduce VAREX (VARied-schema EXtraction), a benchmark for evaluating multimodal foundation models on structured data extraction from government

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

Variational Feature Compression for Model-Specific Representations

DGX agent

arXiv:2604.06644v1 Announce Type: cross Abstract: As deep learning inference is increasingly deployed in shared and cloud-based settings, a growing concern is input repurposing, in which data submitte

model-releasesarxiv-cs-lg
10 Apr 2026
Model Releases

VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics

DGX agent

arXiv:2604.06182v1 Announce Type: cross Abstract: Existing online benchmarks for mobile GUI agents remain largely app-centric and task-homogeneous, failing to reflect the diversity and instability of

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing

DGX agent

arXiv:2604.08401v1 Announce Type: cross Abstract: In large language model (LLM) agents, reasoning trajectories are treated as reliable internal beliefs for guiding actions and updating memory. However

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

VertAX: a differentiable vertex model for learning epithelial tissue mechanics

DGX agent

arXiv:2604.06896v1 Announce Type: new Abstract: Epithelial tissues dynamically reshape through local mechanical interactions among cells, a process well captured by vertex models. Yet their many tunab

model-releasesarxiv-cs-lg
10 Apr 2026
Model Releases

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs

DGX agent

arXiv:2509.08016v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) face a critical bottleneck: increasing the number of input frames to capture fine-grained temporal detail le

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

VisCoder2: Building Multi-Language Visualization Coding Agents

DGX agent

arXiv:2510.23642v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently enabled coding agents capable of generating, executing, and revising visualization code. However, e

model-releasesarxiv-cs-ai
10 Apr 2026
Applications

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment

DGX agent

arXiv:2604.08212v1 Announce Type: new Abstract: General-purpose vision-language models demonstrate strong performance in everyday domains but struggle with specialized technical fields requiring preci

applicationsarxiv-cs-cv
10 Apr 2026
Model Releases

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models

DGX agent

arXiv:2604.07705v1 Announce Type: new Abstract: Aerial vision-and-language navigation (Aerial VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and autonom

model-releasesarxiv-cs-ro
10 Apr 2026
Agents

VisionClaw: Always-On AI Agents through Smart Glasses

DGX agent

arXiv:2604.03486v2 Announce Type: replace-cross Abstract: We present VisionClaw, an always-on wearable AI agent that integrates live egocentric perception with agentic task execution. Running on Meta

agentsarxiv-cs-ai
10 Apr 2026
Model Releases

Visual prompting reimagined: The power of the Activation Prompts

DGX agent

arXiv:2604.06440v1 Announce Type: cross Abstract: Visual prompting (VP) has emerged as a popular method to repurpose pretrained vision models for adaptation to downstream tasks. Unlike conventional mo

model-releasesarxiv-cs-lg
10 Apr 2026
Model Releases

Visually-grounded Humanoid Agents

DGX agent

arXiv:2604.08509v1 Announce Type: new Abstract: Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively

model-releasesarxiv-cs-cv
10 Apr 2026
Safety

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

DGX agent

arXiv:2604.08168v1 Announce Type: new Abstract: Vision-language-action (VLA) models have advanced robot manipulation through large-scale pretraining, but real-world deployment remains challenging due

safetyarxiv-cs-ro
10 Apr 2026
Safety

VLMShield: Efficient and Robust Defense of Vision-Language Models against Malicious Prompts

DGX agent

arXiv:2604.06502v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face significant safety vulnerabilities from malicious prompt attacks due to weakened alignment during visual integration.

safetyarxiv-cs-lg
10 Apr 2026
Model Releases

VSAS-BENCH: Real-Time Evaluation of Visual Streaming Assistant Models

DGX agent

arXiv:2604.07634v1 Announce Type: new Abstract: Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core

model-releasesarxiv-cs-cv
10 Apr 2026
Model Releases

WASD: Locating Critical Neurons as Sufficient Conditions for Explaining and Controlling LLM Behavior

DGX agent

arXiv:2603.18474v2 Announce Type: replace Abstract: Precise behavioral control of large language models (LLMs) is critical for complex applications. However, existing methods often incur high training

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Weakly Supervised Distillation of Hallucination Signals into Transformer Representations

DGX agent

arXiv:2604.06277v1 Announce Type: new Abstract: Existing hallucination detection methods for large language models (LLMs) rely on external verification at inference time, requiring gold answers, retri

model-releasesarxiv-cs-ai
10 Apr 2026
Research

Weakly-Supervised Lung Nodule Segmentation via Training-Free Guidance of 3D Rectified Flow

DGX agent

arXiv:2604.08313v1 Announce Type: new Abstract: Dense annotations, such as segmentation masks, are expensive and time-consuming to obtain, especially for 3D medical images where expert voxel-wise labe

researcharxiv-cs-cv
10 Apr 2026
Research

Weaves, Wires, and Morphisms: Formalizing and Implementing the Algebra of Deep Learning

DGX agent

arXiv:2604.07242v1 Announce Type: new Abstract: Despite deep learning models running well-defined mathematical functions, we lack a formal mathematical framework for describing model architectures. Ad

researcharxiv-cs-lg
10 Apr 2026
Safety

WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search

DGX agent

arXiv:2604.06177v1 Announce Type: cross Abstract: Specialized web tasks in finance, biomedicine, and pharmaceuticals remain challenging due to missing domain priors: queries drift, evidence is noisy,

safetyarxiv-cs-ai
10 Apr 2026
Safety

WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks

DGX agent

arXiv:2604.06367v1 Announce Type: cross Abstract: Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries. While current benchmarks evaluate

safetyarxiv-cs-ai
10 Apr 2026
Research

Weight Group-wise Post-Training Quantization for Medical Foundation Model

DGX agent

arXiv:2604.07674v1 Announce Type: new Abstract: Foundation models have achieved remarkable results in medical image analysis. However, its large network architecture and high computational complexity

researcharxiv-cs-cv
10 Apr 2026
Model Releases

Weighted Bayesian Conformal Prediction

DGX agent

arXiv:2604.06464v1 Announce Type: new Abstract: Conformal prediction provides distribution-free prediction intervals with finite-sample coverage guarantees, and recent work by Snell & Griffiths refra

model-releasesarxiv-cs-lg
10 Apr 2026
Tutorials

What do Language Models Learn and When? The Implicit Curriculum Hypothesis

DGX agent

arXiv:2604.08510v1 Announce Type: new Abstract: Large language models (LLMs) can perform remarkably complex tasks, yet the fine-grained details of how these capabilities emerge during pretraining rema

tutorialsarxiv-cs-cl
10 Apr 2026
Safety

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

DGX agent

arXiv:2604.08524v1 Announce Type: cross Abstract: Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explan

safetyarxiv-cs-cl
10 Apr 2026
Safety

What Makes an Ideal Quote? Recommending 'Unexpected yet Rational' Quotations via Novelty

DGX agent

arXiv:2602.22220v2 Announce Type: replace-cross Abstract: Quotation recommendation aims to enrich writing by suggesting quotes that complement a given context, yet existing systems mostly optimize sur

safetyarxiv-cs-ai
10 Apr 2026
Safety

What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric

DGX agent

arXiv:2604.08494v1 Announce Type: cross Abstract: Scanpath similarity metrics are central to eye-movement research, yet existing methods predominantly evaluate spatial and temporal alignment while neg

safetyarxiv-cs-cl
10 Apr 2026
Model Releases

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning

DGX agent

arXiv:2604.06995v1 Announce Type: new Abstract: Existing Graphical User Interface (GUI) reasoning tasks remain challenging, particularly in UI understanding. Current methods typically rely on direct s

model-releasesarxiv-cs-ai
10 Apr 2026
Research

When Does Context Help? A Systematic Study of Target-Conditional Molecular Property Prediction

DGX agent

arXiv:2604.06558v1 Announce Type: new Abstract: We present the first systematic study of when target context helps molecular property prediction, evaluating context conditioning across 10 diverse prot

researcharxiv-cs-lg
10 Apr 2026
Research

When Fine-Tuning Changes the Evidence: Architecture-Dependent Semantic Drift in Chest X-Ray Explanations

DGX agent

arXiv:2604.08513v1 Announce Type: new Abstract: Transfer learning followed by fine-tuning is widely adopted in medical image classification due to consistent gains in diagnostic performance. However,

researcharxiv-cs-cv
10 Apr 2026
Safety

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

DGX agent

arXiv:2604.08546v1 Announce Type: new Abstract: Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a

safetyarxiv-cs-cv
10 Apr 2026
Model Releases

When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection

DGX agent

arXiv:2510.12476v2 Announce Type: replace Abstract: Large language models (LLMs) have grown more powerful in language generation, producing fluent text and even imitating personal style. Yet, this abi

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't

DGX agent

arXiv:2604.06422v1 Announce Type: cross Abstract: Understanding when Vision-Language Models (VLMs) will behave unexpectedly, whether models can reliably predict their own behavior, and if models adher

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning

DGX agent

arXiv:2604.08281v1 Announce Type: new Abstract: Large reasoning models (LRMs) have achieved strong performance enhancement through scaling test time computation, but due to the inherent limitations of

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models

DGX agent

arXiv:2510.26241v5 Announce Type: replace-cross Abstract: Modern vision-language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Why Are We Lonely? Leveraging LLMs to Measure and Understand Loneliness in Caregivers and Non-caregivers

DGX agent

arXiv:2604.07834v1 Announce Type: new Abstract: This paper presents an LLM-driven approach for constructing diverse social media datasets to measure and compare loneliness in the caregiver and non-car

model-releasesarxiv-cs-cl
10 Apr 2026
← Previous
1…1222122312241225
Next →