AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,376
  • Agents7,554
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,170
  • Local Ai4,930
  • Model Releases23,883
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,376
  • Agents7,554
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,170
  • Local Ai4,930
  • Model Releases23,883
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
Human
88,376Total entries
1Added by human
88,375Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
62,897 results
5 Jun 2026

The Granularity Gap: A Multi-Dimensional Longitudinal Audit of Sycophancy in Gemini Models

Model ReleasesDGX agent

arXiv:2606.05183v1 Announce Type: new Abstract: Large language models are increasingly deployed as high-stakes advisors, yet standard alignment benchmarks treat sycophancy as a binary failure mode. We

The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show

ResearchDGX agent

arXiv:2606.05328v1 Announce Type: cross Abstract: Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators. Yet

The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?

Model ReleasesDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2504.10020v4 Announce Type: replace Abstract: Contrastive decoding strategies are widely used to reduce object hallucinations in multimodal large language models (MLLMs). These methods work by c

The Prosody of Emojis

ApplicationsDGX agent

arXiv:2508.00537v2 Announce Type: replace Abstract: Prosodic features such as pitch, timing, and intonation are central to spoken communication, conveying emotion, intent, and discourse structure. In

The Self-Correction Illusion: LLMs Correct Others but Not Themselves

AgentsDGX agent

arXiv:2606.05976v1 Announce Type: cross Abstract: Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical cl

The Tell-Tale Norm: ell_2 Magnitude as a Signal for Reasoning Dynamics in Large Language Models

ResearchDGX agent

arXiv:2606.06188v1 Announce Type: new Abstract: Recent work has sought to understand Large Language Models (LLMs) reasoning, yet a principled, model-intrinsic signal that captures its layer-wise reaso

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

Model ReleasesDGX agent

arXiv:2606.06476v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have shown strong visual reasoning capabilities, their spatial reasoning abilities remain largely constrained to the

Three-Dimensional Retinal Microvasculature Restoration in OCT Angiography

ResearchDGX agent

arXiv:2606.05375v1 Announce Type: new Abstract: Optical coherence tomographic angiography (OCTA) is a powerful technique for imaging retinal microvasculature. However, acquiring reliable quantificatio

To Be Multimodal or Not to Be: Query-Adaptive Audio-Visual Person Retrieval via Active Modality Detection

ApplicationsDGX agent

arXiv:2606.05931v1 Announce Type: new Abstract: When retrieving a person from a video archive by voice and face, should the system be multimodal or not? In real-world broadcast archives, unlike curate

TopoPult-SSL: Gland-Mask-Free Cross-Device Meibomian Gland Segmentation via Self-Distilled Weak Clinical Priors

Model ReleasesDGX agent

arXiv:2606.05347v1 Announce Type: new Abstract: Every new clinical imaging device creates a domain shift where dense gland masks are expensive yet cheap clinical signals -- eyelid outlines, Pult grade

Toward Culturally Aligned LLMs through Ontology-Guided Multi-Agent Reasoning

SafetyDGX agent

arXiv:2601.21700v3 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly support culturally sensitive decision making, yet often exhibit misalignment due to skewed pretraining dat

Towards a Data Flywheel for Embodied Intelligence in Logistics

SafetyDGX agent

arXiv:2606.05960v1 Announce Type: new Abstract: Embodied intelligence is moving from laboratory demonstrations toward industrial deployment, with the logistics industry serving as a key application sc

Towards Accurate Heart Rate Measurement from Ultra-Short Video Clips via Periodicity-Guided rPPG Estimation and Signal Reconstruction

Model ReleasesDGX agent

arXiv:2506.22078v2 Announce Type: replace Abstract: Many remote Heart Rate (HR) measurement methods focus on estimating remote photoplethysmography (rPPG) signals from video clips lasting around 10 se

Towards Label-Noise Resistant Learning via Optimal Brain Damage Masking

ApplicationsDGX agent

arXiv:2508.09697v3 Announce Type: replace-cross Abstract: Noisy labels are inevitable in real-world scenarios. Due to the strong capacity of deep neural networks to memorize corrupted labels, these no

Towards One-to-Many Temporal Grounding

Model ReleasesDGX agent

arXiv:2606.06294v1 Announce Type: new Abstract: Temporal Grounding (TG) aims to localize video segments corresponding to a textual query. Prior research predominantly focuses on single-segment retriev

Towards Realistic 3D Sonar Simulation

HardwareDGX agent

arXiv:2606.06130v1 Announce Type: new Abstract: As underwater robotics research increasingly addresses complex 3D perception and autonomous navigation, the fidelity of sonar simulation has become a ke

Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs

ResearchDGX agent

arXiv:2606.05846v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) has become a key technology for human--AI interaction. However, code-switching ASR (CS-ASR) remains particularly chal

Trajectory Dynamics in Language Model Hidden States Predict Human Processing Costs Beyond Surprisal

Local AiDGX agent

arXiv:2606.05346v1 Announce Type: new Abstract: Human language comprehension unfolds sequentially: each word is processed in the context of those that came before, and the interpretation builds increm

Two-Way Is Better Than One: Bidirectional Alignment with Cycle Consistency for Exemplar-Free Class-Incremental Learning

SafetyDGX agent

arXiv:2606.05675v1 Announce Type: cross Abstract: Continual learning (CL) seeks models that acquire new skills without erasing prior knowledge. In exemplar-free class-incremental learning (EFCIL), thi

UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning

Model ReleasesDGX agent

arXiv:2606.05576v1 Announce Type: new Abstract: Vision-language models (VLMs) excel on visual question answering and multimodal reasoning benchmarks. Yet their capability on ultra-resolution images -

Uncertainty-Aware Adaptive Sensor Fusion for Autonomous Navigation

HardwareDGX agent

arXiv:2606.05437v1 Announce Type: cross Abstract: This work introduces a hybrid deep learning approach integrated with an Unscented Kalman Filter (UKF) to enhance pose estimation accuracy in Visual-In

UnHype: CLIP-Guided Hypernetworks for Dynamic LoRA Unlearning

ResearchDGX agent

arXiv:2602.03410v2 Announce Type: replace Abstract: Recent advances in large-scale diffusion models have intensified concerns about their potential misuse, particularly in generating realistic yet har

Unifying Dataset Pruning and Distillation for Efficient Large-scale Compression

Model ReleasesDGX agent

arXiv:2502.06434v2 Announce Type: replace Abstract: Dataset pruning (DP) and dataset distillation (DD) fundamentally differ in their outputs: DP selects original image subsets, while DD generates synt

UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matching

Model ReleasesDGX agent

arXiv:2606.05399v1 Announce Type: new Abstract: Existing feed-forward networks excel at predicting a single set of physical properties from visual appearance, but this point-estimate paradigm fundamen

UNIVID: Unified Vision-Language Model for Video Moderation

SafetyDGX agent

arXiv:2606.05748v1 Announce Type: cross Abstract: Global-scale video moderation faces a dual challenge: the need for fine-grained multi-modal reasoning and the demand for interpretable outputs to supp

Unpaired RGB-Thermal Gaussian-Splatting Using Visual Geometric Transformers

SafetyDGX agent

arXiv:2606.05491v1 Announce Type: new Abstract: Multi-modal novel view synthesis (NVS) combining RGB and thermal imagery enables precise 3D scene reconstruction with visual and thermal information. Ho

Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors

ResearchDGX agent

arXiv:2507.12336v2 Announce Type: replace Abstract: Most existing 3D keypoint estimation methods rely on manual annotations or calibrated multi-view images, both of which are expensive to collect. Thi

Unsupervised Skill Discovery for Agentic Data Analysis

AgentsDGX agent

arXiv:2606.06416v1 Announce Type: cross Abstract: Inference-time skill augmentation provides a lightweight way to improve data-analytic agents by injecting reusable procedural knowledge without updati

Unveiling the Unknown: Open Vocabulary Object Detection with Scene Graphs

SafetyDGX agent

arXiv:2606.05916v1 Announce Type: new Abstract: Open-vocabulary object detection seeks to identify novel object categories that were not part of the training data. Many knowledge distillation-based ap

USAD 2.0: Scaling Representation Distillation for Universal Audio Understanding

ResearchDGX agent

arXiv:2606.06444v1 Announce Type: cross Abstract: Audio encoders are critical to modern audio applications as large language models (LLMs) increasingly rely on a single encoder for diverse inputs. Whi

Using Large Language Models to Support High Volume Application Review for an Undergraduate Research Program

Model ReleasesDGX agent

arXiv:2606.05564v1 Announce Type: new Abstract: Undergraduate research programs such as the Summer Undergraduate Research Fellowship (SURF) at Purdue University receive thousands of applications every

Using street view images and visual LLMs to predict heritage values for governance support: Risks, ethics, and policy implications

SafetyDGX agent

arXiv:2601.06056v2 Announce Type: replace-cross Abstract: During 2025 and 2026, the Energy Performance of Buildings Directive is being implemented in the European Union member states, requiring all me

V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation

Model ReleasesDGX agent

arXiv:2606.05665v1 Announce Type: new Abstract: Video-to-video (V2V) generation is difficult to evaluate because outputs must both follow editing instructions and preserve frame-level correspondence w

Value-and-Structure Alignment for Routing-Consistent Quantization of Mixture-of-Experts Models

SafetyDGX agent

arXiv:2606.05688v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale foundation models efficiently by activating only a subset of experts for each token, but their large number of exp

VASO: Formally Verifiable Self-Evolving Skills for Physical AI Agents

Local AiDGX agent

arXiv:2606.05395v1 Announce Type: new Abstract: Reusable robot skills are becoming the basic units through which embodied agents turn open-ended instructions into long-horizon physical behavior. We ar

Vavanagi: a Community-run Platform for Documentation of the Hula Language in Papua New Guinea

ResearchDGX agent

arXiv:2603.14210v2 Announce Type: replace Abstract: We present Vavanagi, a community-run platform for Hula (Vula'a), an Austronesian language of Papua New Guinea with approximately 10,000 speakers. Va

ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation

SafetyDGX agent

arXiv:2606.05718v1 Announce Type: new Abstract: On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher. In m

Video-Rate Streaming Stylization on a Vision-Aware MLLM-Conditioned Edit Diffusion: Asymmetric Batched Inference on a Distilled UNet + MLLM Text Encoder

Model ReleasesDGX agent

arXiv:2606.05981v1 Announce Type: new Abstract: Aggressive distillation of the diffusion U-Net inverts the per-frame bottleneck of real-time text-to-image pipelines: once the denoiser is a 4-step or 1

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

Model ReleasesDGX agent

arXiv:2606.05259v1 Announce Type: new Abstract: We introduce VideoKR, the first large-scale training corpus specifically designed to strengthen knowledge- and reasoning-intensive video understanding.

Vision Hopfield Memory Networks

Local AiDGX agent

arXiv:2603.25157v2 Announce Type: replace-cross Abstract: Recent vision and multimodal foundation backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable pr

Visual Commonsense Driven Knowledge Refinements for Scene Graph Generation

ResearchDGX agent

arXiv:2606.06369v1 Announce Type: new Abstract: Learning-driven Scene Graph Generation (SGG) models excel on frequent relation types but degrade sharply under annotation sparsity, failing to capture r

Visuotactile and Explicitly Force-Controlled Robotic Ultrasound for Abdominal Volumetric Reconstruction

AgentsDGX agent

arXiv:2606.05848v1 Announce Type: new Abstract: In this paper, we present a robotic ultrasound acquisition system that integrates stereo vision, touch-based feedback, and expert-informed strategies to

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation

SafetyDGX agent

arXiv:2510.23497v3 Announce Type: replace Abstract: Training vision-language models (VLMs) for complex reasoning remains a challenging task, i.a. due to the scarcity of high-quality image-text reasoni

VOLT: Vision and Language Trajectory Segmentation for Faster-than-Demonstration Policies

SafetyDGX agent

arXiv:2606.06323v1 Announce Type: new Abstract: Humans often take longer to demonstrate a task than a robot would need to execute it. Rather than learning to replicate the demonstration at the same pa

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning

Model ReleasesDGX agent

arXiv:2606.05736v1 Announce Type: new Abstract: Video reasoning aims to understand complex temporal events and causal relationships within videos. Recently, Chain-of-Thought (CoT) has been introduced

VZCrash: A Large-Scale IMU Dataset of Ego-Vehicle Crashes

Model ReleasesDGX agent

arXiv:2606.06074v1 Announce Type: new Abstract: We introduce VZCrash, the largest publicly available dataset of real-world vehicle collision data featuring Inertial Measurement Unit (IMU) telemetry. T

Wave Focusing in Metamaterials: Tactile Displays Beyond the Diffraction Limit

ResearchDGX agent

arXiv:2606.05572v1 Announce Type: cross Abstract: We address the challenge of engineering distributed haptic displays capable of reproducing multiple localized, independently addressable vibrations --

Waypoints Matter: A Systematic Study for Sampling-Based Trajectory Planning

Model ReleasesDGX agent

arXiv:2606.06366v1 Announce Type: new Abstract: Real-time autonomous driving commonly relies on sampling-based trajectory planners that link candidate trajectories to target waypoints along the road c

What Makes Two Language Models Think Alike?

ResearchDGX agent

arXiv:2406.12620v3 Announce Type: replace Abstract: Do architectural and training differences influence the way models represent and process language? Traditional similarity metrics tell us whether tw

What Objects Enable, Not What They Are: Functional Latent Spaces for Affordance Reasoning

ResearchDGX agent

arXiv:2606.05533v1 Announce Type: cross Abstract: Existing robot planning systems rely on appearance-based reasoning, where visual observations are encoded into latent spaces organized around object a

What's in a Name? Morphological Shortcuts by LLMs in Pharmacology

SafetyDGX agent

arXiv:2606.05616v1 Announce Type: new Abstract: The morphological form of a word can often give cues to its meaning, but purely relying on these mappings can lead to overgeneralization in high-stakes

What's Under the Skin? Estimating Swine Body Condition

ApplicationsDGX agent

arXiv:2606.05611v1 Announce Type: new Abstract: Sow body condition is an important indicator for growers as it has a large impact on lactation performance and piglet survival. However, body condition

When AI Says It Feels

SafetyDGX agent

arXiv:2606.05734v1 Announce Type: cross Abstract: Large language models (LLMs) are generally constrained from expressing feelings through human-preference alignment in post-training processes. This po

When Evidence is Sparse: Weakly Supervised Early Failure Alerting in Dialogs and LLM-Agent Trajectories

SafetyDGX agent

arXiv:2606.05414v1 Announce Type: new Abstract: Early failure alerting requires deciding, while a dialog or agent trajectory is still unfolding, whether to flag it as likely to fail. This is challengi

When New Generators Arrive: Lifelong Machine-Generated Text Attribution via Ridge Feature Transfer

ResearchDGX agent

arXiv:2606.05626v1 Announce Type: new Abstract: Machine-generated text (MGT) attribution aims to identify the specific generator responsible for a given text, thereby providing fine-grained evidence f

Where does Absolute Position come from in decoder-only Transformers?

ResearchDGX agent

arXiv:2606.06160v1 Announce Type: cross Abstract: RoPE-trained transformers distinguish absolute position in their attention patterns, even though RoPE encodes only relative offsets in the inner produ

Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

Local AiDGX agent

arXiv:2606.06113v1 Announce Type: new Abstract: Despite generating increasingly photorealistic images, text-to-image (T2I) models still exhibit localized, subtle, and structurally complex failures. Di

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

HardwareDGX agent

arXiv:2606.05979v1 Announce Type: new Abstract: We propose world-language-action (WLA) models as a new class of embodied foundation models. WLA takes textual instructions, images, and robot states as

Would you still call this Dax? Novel Visual References in VLMs and Humans

Model ReleasesDGX agent

arXiv:2606.05409v1 Announce Type: cross Abstract: Vision-language models (VLMs), like human learners, are frequently exposed to new visual concepts, but how they map novel visual references to languag

You Only Index Once: Cross-Layer Sparse Attention with Shared Routing

ResearchDGX agent

arXiv:2606.06467v1 Announce Type: new Abstract: Long-context inference in modern LLMs is increasingly constrained by decoding efficiency, especially in reasoning-heavy settings where models generate l

← Previous
1…477478479480481…1049
Next →