AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
Human
85,115Total entries
1Added by human
85,114Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
60,292 results
10 Apr 2026

Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator

ResearchDGX agent

arXiv:2604.08121v1 Announce Type: new Abstract: Unified multimodal models integrating visual understanding and generation face a fundamental challenge: visual generation incurs substantially higher co

Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs

Model ReleasesDGX agent

arXiv:2601.21463v2 Announce Type: replace-cross Abstract: Existing speech editing detection (SED) datasets are predominantly constructed using manual splicing or limited editing operations, resulting

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models

ApplicationsDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2602.20231v2 Announce Type: replace-cross Abstract: Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-acti

UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding

Model ReleasesDGX agent

arXiv:2604.08522v1 Announce Type: new Abstract: Video temporal grounding (VTG) is typically tackled with dataset-specific models that transfer poorly across domains and query styles. Recent efforts to

Unsupervised Neural Network for Automated Classification of Surgical Urgency Levels in Medical Transcriptions

ApplicationsDGX agent

arXiv:2604.06214v1 Announce Type: cross Abstract: Efficient classification of surgical procedures by urgency is paramount to optimize patient care and resource allocation within healthcare systems. Th

URMF: Uncertainty-aware Robust Multimodal Fusion for Multimodal Sarcasm Detection

SafetyDGX agent

arXiv:2604.06728v1 Announce Type: cross Abstract: Multimodal sarcasm detection (MSD) aims to identify sarcastic intent from semantic incongruity between text and image. Although recent methods have im

Validated Intent Compilation for Constrained Routing in LEO Mega-Constellations

Model ReleasesDGX agent

arXiv:2604.07264v1 Announce Type: cross Abstract: Operating LEO mega-constellations requires translating high-level operator intents ('reroute financial traffic away from polar links under 80 ms') int

VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents

Model ReleasesDGX agent

arXiv:2603.15118v2 Announce Type: replace Abstract: We introduce VAREX (VARied-schema EXtraction), a benchmark for evaluating multimodal foundation models on structured data extraction from government

Variational Feature Compression for Model-Specific Representations

Model ReleasesDGX agent

arXiv:2604.06644v1 Announce Type: cross Abstract: As deep learning inference is increasingly deployed in shared and cloud-based settings, a growing concern is input repurposing, in which data submitte

VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics

Model ReleasesDGX agent

arXiv:2604.06182v1 Announce Type: cross Abstract: Existing online benchmarks for mobile GUI agents remain largely app-centric and task-homogeneous, failing to reflect the diversity and instability of

Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing

Model ReleasesDGX agent

arXiv:2604.08401v1 Announce Type: cross Abstract: In large language model (LLM) agents, reasoning trajectories are treated as reliable internal beliefs for guiding actions and updating memory. However

VertAX: a differentiable vertex model for learning epithelial tissue mechanics

Model ReleasesDGX agent

arXiv:2604.06896v1 Announce Type: new Abstract: Epithelial tissues dynamically reshape through local mechanical interactions among cells, a process well captured by vertex models. Yet their many tunab

Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs

Model ReleasesDGX agent

arXiv:2509.08016v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) face a critical bottleneck: increasing the number of input frames to capture fine-grained temporal detail le

VisCoder2: Building Multi-Language Visualization Coding Agents

Model ReleasesDGX agent

arXiv:2510.23642v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently enabled coding agents capable of generating, executing, and revising visualization code. However, e

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment

ApplicationsDGX agent

arXiv:2604.08212v1 Announce Type: new Abstract: General-purpose vision-language models demonstrate strong performance in everyday domains but struggle with specialized technical fields requiring preci

Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models

Model ReleasesDGX agent

arXiv:2604.07705v1 Announce Type: new Abstract: Aerial vision-and-language navigation (Aerial VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and autonom

VisionClaw: Always-On AI Agents through Smart Glasses

AgentsDGX agent

arXiv:2604.03486v2 Announce Type: replace-cross Abstract: We present VisionClaw, an always-on wearable AI agent that integrates live egocentric perception with agentic task execution. Running on Meta

Visual prompting reimagined: The power of the Activation Prompts

Model ReleasesDGX agent

arXiv:2604.06440v1 Announce Type: cross Abstract: Visual prompting (VP) has emerged as a popular method to repurpose pretrained vision models for adaptation to downstream tasks. Unlike conventional mo

Visually-grounded Humanoid Agents

Model ReleasesDGX agent

arXiv:2604.08509v1 Announce Type: new Abstract: Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

SafetyDGX agent

arXiv:2604.08168v1 Announce Type: new Abstract: Vision-language-action (VLA) models have advanced robot manipulation through large-scale pretraining, but real-world deployment remains challenging due

VLMShield: Efficient and Robust Defense of Vision-Language Models against Malicious Prompts

SafetyDGX agent

arXiv:2604.06502v1 Announce Type: new Abstract: Vision-Language Models (VLMs) face significant safety vulnerabilities from malicious prompt attacks due to weakened alignment during visual integration.

VSAS-BENCH: Real-Time Evaluation of Visual Streaming Assistant Models

Model ReleasesDGX agent

arXiv:2604.07634v1 Announce Type: new Abstract: Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core

WASD: Locating Critical Neurons as Sufficient Conditions for Explaining and Controlling LLM Behavior

Model ReleasesDGX agent

arXiv:2603.18474v2 Announce Type: replace Abstract: Precise behavioral control of large language models (LLMs) is critical for complex applications. However, existing methods often incur high training

Weakly Supervised Distillation of Hallucination Signals into Transformer Representations

Model ReleasesDGX agent

arXiv:2604.06277v1 Announce Type: new Abstract: Existing hallucination detection methods for large language models (LLMs) rely on external verification at inference time, requiring gold answers, retri

Weakly-Supervised Lung Nodule Segmentation via Training-Free Guidance of 3D Rectified Flow

ResearchDGX agent

arXiv:2604.08313v1 Announce Type: new Abstract: Dense annotations, such as segmentation masks, are expensive and time-consuming to obtain, especially for 3D medical images where expert voxel-wise labe

Weaves, Wires, and Morphisms: Formalizing and Implementing the Algebra of Deep Learning

ResearchDGX agent

arXiv:2604.07242v1 Announce Type: new Abstract: Despite deep learning models running well-defined mathematical functions, we lack a formal mathematical framework for describing model architectures. Ad

WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search

SafetyDGX agent

arXiv:2604.06177v1 Announce Type: cross Abstract: Specialized web tasks in finance, biomedicine, and pharmaceuticals remain challenging due to missing domain priors: queries drift, evidence is noisy,

WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks

SafetyDGX agent

arXiv:2604.06367v1 Announce Type: cross Abstract: Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries. While current benchmarks evaluate

Weight Group-wise Post-Training Quantization for Medical Foundation Model

ResearchDGX agent

arXiv:2604.07674v1 Announce Type: new Abstract: Foundation models have achieved remarkable results in medical image analysis. However, its large network architecture and high computational complexity

Weighted Bayesian Conformal Prediction

Model ReleasesDGX agent

arXiv:2604.06464v1 Announce Type: new Abstract: Conformal prediction provides distribution-free prediction intervals with finite-sample coverage guarantees, and recent work by Snell & Griffiths refra

What do Language Models Learn and When? The Implicit Curriculum Hypothesis

TutorialsDGX agent

arXiv:2604.08510v1 Announce Type: new Abstract: Large language models (LLMs) can perform remarkably complex tasks, yet the fine-grained details of how these capabilities emerge during pretraining rema

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

SafetyDGX agent

arXiv:2604.08524v1 Announce Type: cross Abstract: Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explan

What Makes an Ideal Quote? Recommending 'Unexpected yet Rational' Quotations via Novelty

SafetyDGX agent

arXiv:2602.22220v2 Announce Type: replace-cross Abstract: Quotation recommendation aims to enrich writing by suggesting quotes that complement a given context, yet existing systems mostly optimize sur

What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metric

SafetyDGX agent

arXiv:2604.08494v1 Announce Type: cross Abstract: Scanpath similarity metrics are central to eye-movement research, yet existing methods predominantly evaluate spatial and temporal alignment while neg

What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI Reasoning

Model ReleasesDGX agent

arXiv:2604.06995v1 Announce Type: new Abstract: Existing Graphical User Interface (GUI) reasoning tasks remain challenging, particularly in UI understanding. Current methods typically rely on direct s

When Does Context Help? A Systematic Study of Target-Conditional Molecular Property Prediction

ResearchDGX agent

arXiv:2604.06558v1 Announce Type: new Abstract: We present the first systematic study of when target context helps molecular property prediction, evaluating context conditioning across 10 diverse prot

When Fine-Tuning Changes the Evidence: Architecture-Dependent Semantic Drift in Chest X-Ray Explanations

ResearchDGX agent

arXiv:2604.08513v1 Announce Type: new Abstract: Transfer learning followed by fine-tuning is widely adopted in medical image classification due to consistent gains in diagnostic performance. However,

When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

SafetyDGX agent

arXiv:2604.08546v1 Announce Type: new Abstract: Text-to-video diffusion models have enabled open-ended video synthesis, but often struggle with generating the correct number of objects specified in a

When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection

Model ReleasesDGX agent

arXiv:2510.12476v2 Announce Type: replace Abstract: Large language models (LLMs) have grown more powerful in language generation, producing fluent text and even imitating personal style. Yet, this abi

When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't

Model ReleasesDGX agent

arXiv:2604.06422v1 Announce Type: cross Abstract: Understanding when Vision-Language Models (VLMs) will behave unexpectedly, whether models can reliably predict their own behavior, and if models adher

When to Trust Tools? Adaptive Tool Trust Calibration For Tool-Integrated Math Reasoning

Model ReleasesDGX agent

arXiv:2604.08281v1 Announce Type: new Abstract: Large reasoning models (LRMs) have achieved strong performance enhancement through scaling test time computation, but due to the inherent limitations of

Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models

Model ReleasesDGX agent

arXiv:2510.26241v5 Announce Type: replace-cross Abstract: Modern vision-language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not

Why Are We Lonely? Leveraging LLMs to Measure and Understand Loneliness in Caregivers and Non-caregivers

Model ReleasesDGX agent

arXiv:2604.07834v1 Announce Type: new Abstract: This paper presents an LLM-driven approach for constructing diverse social media datasets to measure and compare loneliness in the caregiver and non-car

'Why This Avoidance Maneuver?' Contrastive Explanations in Human-Supervised Maritime Autonomous Navigation

AgentsDGX agent

arXiv:2604.08032v1 Announce Type: cross Abstract: Automated maritime collision avoidance will rely on human supervision for the foreseeable future. This necessitates transparency into how the system p

Working Paper: Towards a Category-theoretic Comparative Framework for Artificial General Intelligence

AgentsDGX agent

arXiv:2603.28906v2 Announce Type: replace Abstract: AGI has become the Holly Grail of AI with the promise of level intelligence and the major Tech companies around the world are investing unprecedente

WorldMAP: Bootstrapping Vision-Language Navigation Trajectory Prediction with Generative World Models

ResearchDGX agent

arXiv:2604.07957v1 Announce Type: cross Abstract: Vision-language models (VLMs) and generative world models are opening new opportunities for embodied navigation. VLMs are increasingly used as direct

WRAP++: Web discoveRy Amplified Pretraining

ResearchDGX agent

arXiv:2604.06829v2 Announce Type: cross Abstract: Synthetic data rephrasing has emerged as a powerful technique for enhancing knowledge acquisition during large language model (LLM) pretraining. Howev

WUTDet: A 100K-Scale Ship Detection Dataset and Benchmarks with Dense Small Objects

ResearchDGX agent

arXiv:2604.07759v1 Announce Type: new Abstract: Ship detection for navigation is a fundamental perception task in intelligent waterway transportation systems. However, existing public ship detection d

XR-CareerAssist: An Immersive Platform for Personalised Career Guidance Leveraging Extended Reality and Multimodal AI

ResearchDGX agent

arXiv:2604.06901v1 Announce Type: cross Abstract: Conventional career guidance platforms rely on static, text-driven interfaces that struggle to engage users or deliver personalised, evidence-based in

You Point, I Learn: Online Adaptation of Interactive Segmentation Models for Handling Distribution Shifts in Medical Imaging

Model ReleasesDGX agent

arXiv:2503.06717v3 Announce Type: replace Abstract: Interactive segmentation uses real-time user inputs, such as mouse clicks, to iteratively refine model predictions. Although not originally designed

Zatom-1: A Multimodal Flow Foundation Model for 3D Molecules and Materials

ResearchDGX agent

arXiv:2602.22251v3 Announce Type: replace-cross Abstract: General-purpose 3D chemical modeling encompasses molecules and materials, requiring both generative and predictive capabilities. However, most

ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training

HardwareDGX agent

arXiv:2603.04385v3 Announce Type: replace Abstract: Feed-forward transformer models have driven rapid progress in 3D vision, but state-of-the-art methods such as VGGT and pi^3 have a computational

← Previous
1…100310041005
Next →