AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,271
  • Agents7,543
  • Applications5,408
  • Concepts5
  • Hardware1,828
  • Industry6,162
  • Local Ai4,927
  • Model Releases23,818
  • Research20,122
  • Safety13,367
  • Syntheses17
  • Tools1,674
  • Tutorials3,400

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,271
  • Agents7,543
  • Applications5,408
  • Concepts5
  • Hardware1,828
  • Industry6,162
  • Local Ai4,927
  • Model Releases23,818
  • Research20,122
  • Safety13,367
  • Syntheses17
  • Tools1,674
  • Tutorials3,400

Source
Human
88,271Total entries
1Added by human
88,270Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
62,897 results
Model Releases

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

DGX agent

arXiv:2510.08759v2 Announce Type: replace Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, exi

model-releasesarxiv-cs-cv
22 May 2026
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Distributed Image Compression with Multimodal Side Information at Extremely Low Bitrates

DGX agent

arXiv:2605.22061v1 Announce Type: new Abstract: Distributed Image Compression (DIC) is crucial for multi-view transmission, especially when operating at extremely low bitrates (< 0.1 bpp). Its core ch

safetyarxiv-cs-cv
22 May 2026
Agents

Distributed Multi-Coverage for Robot Swarms

DGX agent

arXiv:2605.21686v1 Announce Type: new Abstract: Autonomous drone swarms deployed for surveillance, environmental monitoring, and infrastructure inspection must maintain reliable coverage of critical a

agentsarxiv-cs-ro
22 May 2026
Tutorials

Diverge to Induce Prompting: Multi-Rationale Induction for Zero-Shot Reasoning

DGX agent

arXiv:2602.08028v1 Announce Type: cross Abstract: To address the instability of unguided reasoning paths in standard Chain-of-Thought prompting, recent methods guide large language models (LLMs) by fi

tutorialsarxiv-cs-ai
22 May 2026
Model Releases

Diverse Yet Consistent: Context-Guided Diffusion with Energy-Based Joint Refinement for Multi-Agent Motion Prediction

DGX agent

arXiv:2605.22017v1 Announce Type: new Abstract: Deepgenerative models havebecomeapromisingapproach for human motion prediction due to their ability to capture multimodal distributions and represent di

model-releasesarxiv-cs-cv
22 May 2026
Research

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?

DGX agent

arXiv:2605.22170v1 Announce Type: new Abstract: In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges abo

researcharxiv-cs-cl
22 May 2026
Model Releases

Do Vision Models Encode Object-Level Semantic Relatedness? A Cognitive Psychology-Inspired Benchmark

DGX agent

arXiv:1709.03806v2 Announce Type: replace Abstract: Modern vision models have achieved strong object-recognition performance, yet it remains unclear whether their representations encode object-level s

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Does Slightly Mean Somewhat? Measuring Vague Intensity Words in LLM Numeric Actions

DGX agent

arXiv:2605.21827v1 Announce Type: new Abstract: Do language models preserve the ordinal meaning of intensity words when those words must produce numeric actions? I study a researcher-constructed scale

model-releasesarxiv-cs-cl
22 May 2026
Hardware

Don't Collapse Your Features: Why CenterLoss Hurts OOD Detection and Multi-Scale Mahalanobis Wins

DGX agent

arXiv:2605.21493v1 Announce Type: cross Abstract: The ability to detect out-of-distribution (OOD) inputs is fundamental to safe deployment of machine learning systems. Yet, current methods often rely

hardwarearxiv-cs-cv
22 May 2026
Model Releases

Dual-Integrated Low-Latency Single-Lens Infrared Computational Imaging for Object Detection

DGX agent

arXiv:2605.21964v1 Announce Type: new Abstract: Computational imaging enables compact infrared systems, but deep-learning pipelines that combine image reconstruction and object detection often introdu

model-releasesarxiv-cs-cv
22 May 2026
Hardware

EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation

DGX agent

arXiv:2605.22051v1 Announce Type: new Abstract: Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of sp

hardwarearxiv-cs-cv
22 May 2026
Agents

Echo: Learning from Experience Data via User-Driven Refinement

DGX agent

arXiv:2605.21984v1 Announce Type: cross Abstract: Static 'human data' faces inherent limitations: it is expensive to scale and bounded by the knowledge of its creators. Continuous learning from 'exper

agentsarxiv-cs-cl
22 May 2026
Safety

Echo4DIR: 4D Implicit Heart Reconstruction from 2D Echocardiography Videos

DGX agent

arXiv:2605.22066v1 Announce Type: new Abstract: Reconstructing 4D (3D+t) cardiac geometry from sparse 2D echocardiography is highly desirable yet fundamentally challenged by geometric ambiguity and te

safetyarxiv-cs-cv
22 May 2026
Model Releases

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

DGX agent

arXiv:2605.22138v1 Announce Type: cross Abstract: How should an agent decide when and how to plan? A dominant approach builds agents as reactive policies with adaptive computation (e.g., chain-of-thou

model-releasesarxiv-cs-cl
22 May 2026
Applications

ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing

DGX agent

arXiv:2605.20802v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) exploit event-driven and addition-only computation to substantially improve efficiency for intelligent computation. A k

applicationsarxiv-cs-ai
22 May 2026
Agents

Enabling Regulatory Multi-Agent Collaboration: Architecture, Challenges, and Solutions

DGX agent

arXiv:2509.09215v2 Announce Type: replace Abstract: Large language models (LLMs)-empowered autonomous agents are transforming both digital and physical environments by enabling adaptive, multi-agent c

agentsarxiv-cs-ai
22 May 2026
Safety

Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention

DGX agent

arXiv:2605.21842v1 Announce Type: cross Abstract: Standard transformer attention computes pairwise similarity between queries and keys, treating all tokens as equally salient regardless of their intri

safetyarxiv-cs-cl
22 May 2026
Research

Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning

DGX agent

arXiv:2401.00139v3 Announce Type: replace-cross Abstract: This paper introduces a causal attribution model to enhance the interpretability of large language models (LLMs) and improve their causal reas

researcharxiv-cs-cl
22 May 2026
Research

Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following

DGX agent

arXiv:2605.22607v1 Announce Type: new Abstract: Gaze following requires both scene understanding and gaze reasoning to localize the gaze target of an in-scene person. Recently, vision foundation model

researcharxiv-cs-cv
22 May 2026
Safety

Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

DGX agent

arXiv:2605.22185v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, thei

safetyarxiv-cs-cv
22 May 2026
Research

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding

DGX agent

arXiv:2605.22078v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video understanding tasks, yet challenges remain in efficientl

researcharxiv-cs-cv
22 May 2026
Research

EntmaxKV: Support-Aware Decoding for Entmax Attention

DGX agent

arXiv:2605.21649v1 Announce Type: cross Abstract: Long-context decoding is increasingly limited by KV-cache memory traffic since each generated token attends over a cache whose size grows linearly wit

researcharxiv-cs-cl
22 May 2026
Research

Entropy-Guided Self-Supervised Learning for Medical Image Classification

DGX agent

arXiv:2605.21970v1 Announce Type: cross Abstract: Accurate and robust medical image classification is paramount for early disease diagnosis and treatment planning. However, challenges such as limited

researcharxiv-cs-cv
22 May 2026
Research

Epicure: Navigating the Emergent Geometry of Food Ingredient Embeddings

DGX agent

arXiv:2605.22391v1 Announce Type: cross Abstract: We present Epicure, a family of three sibling skip-gram ingredient embeddings retrained from scratch on a multilingual recipe corpus. We aggregate 4.1

researcharxiv-cs-cl
22 May 2026
Model Releases

Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark

DGX agent

arXiv:2503.17599v3 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated considerable potential in general practice. However, existing benchmarks and evaluation frameworks pr

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Evaluating Commercial AI Chatbots as News Intermediaries

DGX agent

arXiv:2605.22785v1 Announce Type: new Abstract: AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their p

model-releasesarxiv-cs-cl
22 May 2026
Agents

Evaluating multimodal emotion recognition in proactive conversational agents: A user study

DGX agent

arXiv:2605.20200v1 Announce Type: cross Abstract: This article presents a multimodal emotion recognition module integrated into a proactive Socially Interactive Agent (SIA) powered by generative artif

agentsarxiv-cs-ai
22 May 2026
Model Releases

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

DGX agent

arXiv:2605.20630v1 Announce Type: new Abstract: Industrial asset operations workflows are latency-sensitive because a single user query may require coordination over sensor data, work orders, failure

model-releasesarxiv-cs-ai
22 May 2026
Research

Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents

DGX agent

arXiv:2605.22203v1 Announce Type: new Abstract: In this study, we compare the performance of four text chunking approaches: Recursive, Khmer-Aware, Sentence-Based, and LLM-Based within a Retrieval-Aug

researcharxiv-cs-cl
22 May 2026
Model Releases

Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset

DGX agent

arXiv:2605.22186v1 Announce Type: new Abstract: Event-based low-light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking t

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

EventGait: Towards Robust Gait Recognition with Event Streams

DGX agent

arXiv:2605.22139v1 Announce Type: new Abstract: Gait recognition enables non-intrusive, privacy-preserving identification but suffers in uncontrolled environments due to illumination and motion sensit

model-releasesarxiv-cs-cv
22 May 2026
Agents

EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning

DGX agent

arXiv:2605.22208v1 Announce Type: new Abstract: Multimodal Large Language Model (MLLM)-driven image restoration agent demonstrates effectiveness in degradation coupling scenarios by flexibly selecting

agentsarxiv-cs-cv
22 May 2026
Research

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control

DGX agent

arXiv:2605.21862v1 Announce Type: new Abstract: Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone. Yet r

researcharxiv-cs-ro
22 May 2026
Agents

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models

DGX agent

arXiv:2605.21931v1 Announce Type: new Abstract: Recent Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in video reasoning through reinforcement learning (RL). However, e

agentsarxiv-cs-cv
22 May 2026
Applications

Exposing Vulnerabilities in Visible-Infrared VLMs: A Unified Geometric Adversarial Framework with Cross-Task Transferability

DGX agent

arXiv:2605.22273v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, but their adversarial robustness in visible-infrared (VI

applicationsarxiv-cs-cv
22 May 2026
Research

Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention

DGX agent

arXiv:2605.22072v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for advancing complex reasoning in large language models, and

researcharxiv-cs-cl
22 May 2026
Model Releases

FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

DGX agent

arXiv:2605.22552v1 Announce Type: new Abstract: Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is

model-releasesarxiv-cs-cv
22 May 2026
Local Ai

FastTab: A Fast Table Recognizer with a Tiny Recursive Module and 1D Transformers

DGX agent

arXiv:2605.22422v1 Announce Type: new Abstract: Table structure recognition (TSR) requires both table-level coherence (row/column counts, headers, spanning cells) and precise separator localization. W

local-aiarxiv-cs-cv
22 May 2026
Model Releases

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

DGX agent

arXiv:2605.21625v1 Announce Type: cross Abstract: The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus

model-releasesarxiv-cs-cl
22 May 2026
Research

Flow-based Gaussian Splatting for Continuous-Scale Remote Sensing Image Super-Resolution

DGX agent

arXiv:2605.22147v1 Announce Type: new Abstract: High-resolution remote sensing images (RSIs) are crucial for Earth observation applications, yet acquiring them is often limited by sensor constraints a

researcharxiv-cs-cv
22 May 2026
Agents

Flying Together: Human-Guided Immersive Shared Control for Aerial Robot Teams in Unknown Environments

DGX agent

arXiv:2605.21680v1 Announce Type: new Abstract: While autonomous multi-robots can achieve safe and coordinated navigation, they often struggle to adapt to unforeseen conditions and to capture operator

agentsarxiv-cs-ro
22 May 2026
Safety

FlyRoute: Self-Evolving Agent Profiling via Data Flywheel for Adaptive Task Routing

DGX agent

arXiv:2605.22057v1 Announce Type: new Abstract: Enterprise routers assign queries to expert agents, yet deployed profiles stay static while agents evolve (prompts, tools, models), and developers rarel

safetyarxiv-cs-cl
22 May 2026
Safety

Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain

DGX agent

arXiv:2602.17186v2 Announce Type: replace Abstract: Large Vision Language Models (LVLMs) have achieved remarkable progress, yet they often suffer from language bias, producing answers without relying

safetyarxiv-cs-cv
22 May 2026
Research

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding

DGX agent

arXiv:2605.21973v1 Announce Type: new Abstract: Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream,

researcharxiv-cs-cv
22 May 2026
Research

ForeSplat: Optimization-Aware Foresight for Feed-Forward 3D Gaussian Splatting

DGX agent

arXiv:2605.22020v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) models offer fast single-pass reconstruction,but scaling them to match per-scene optimization quality is funda

researcharxiv-cs-cv
22 May 2026
Model Releases

FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments

DGX agent

arXiv:2605.22018v1 Announce Type: new Abstract: The Flooded Road Environments Dataset (FRED) is, to our knowledge, the first multi-modal autonomous driving dataset specifically targeting the collectio

model-releasesarxiv-cs-cv
22 May 2026
Safety

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

DGX agent

arXiv:2605.22671v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior

safetyarxiv-cs-cv
22 May 2026
Agents

From Automated to Autonomous: Hierarchical Agent-native Network Architecture (HANA)

DGX agent

arXiv:2605.20608v1 Announce Type: new Abstract: Realizing Level 4/5 Autonomous Networks (AN) demands a shift from static automation to agent-native intelligence. Current operations, reliant on rigid s

agentsarxiv-cs-ai
22 May 2026
← Previous
1…792793794795796…1311
Next →