AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

Content type
88,246Total entries
1Added by human
88,245Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
62,897 results
Model Releases

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

DGX agent

arXiv:2510.08759v2 Announce Type: replace Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, exi

model-releasesarxiv-cs-cv
22 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Safety

Distributed Image Compression with Multimodal Side Information at Extremely Low Bitrates

DGX agent

arXiv:2605.22061v1 Announce Type: new Abstract: Distributed Image Compression (DIC) is crucial for multi-view transmission, especially when operating at extremely low bitrates (< 0.1 bpp). Its core ch

safetyarxiv-cs-cv
22 May 2026
Agents

Distributed Multi-Coverage for Robot Swarms

DGX agent

arXiv:2605.21686v1 Announce Type: new Abstract: Autonomous drone swarms deployed for surveillance, environmental monitoring, and infrastructure inspection must maintain reliable coverage of critical a

agentsarxiv-cs-ro
22 May 2026
Tutorials

Diverge to Induce Prompting: Multi-Rationale Induction for Zero-Shot Reasoning

DGX agent

arXiv:2602.08028v1 Announce Type: cross Abstract: To address the instability of unguided reasoning paths in standard Chain-of-Thought prompting, recent methods guide large language models (LLMs) by fi

tutorialsarxiv-cs-ai
22 May 2026
Model Releases

Diverse Yet Consistent: Context-Guided Diffusion with Energy-Based Joint Refinement for Multi-Agent Motion Prediction

DGX agent

arXiv:2605.22017v1 Announce Type: new Abstract: Deepgenerative models havebecomeapromisingapproach for human motion prediction due to their ability to capture multimodal distributions and represent di

model-releasesarxiv-cs-cv
22 May 2026
Research

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?

DGX agent

arXiv:2605.22170v1 Announce Type: new Abstract: In recent years, several Speech Language Models (SLMs) that represent speech and written text jointly have been presented. The question then emerges abo

researcharxiv-cs-cl
22 May 2026
Model Releases

Do Vision Models Encode Object-Level Semantic Relatedness? A Cognitive Psychology-Inspired Benchmark

DGX agent

arXiv:1709.03806v2 Announce Type: replace Abstract: Modern vision models have achieved strong object-recognition performance, yet it remains unclear whether their representations encode object-level s

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Does Slightly Mean Somewhat? Measuring Vague Intensity Words in LLM Numeric Actions

DGX agent

arXiv:2605.21827v1 Announce Type: new Abstract: Do language models preserve the ordinal meaning of intensity words when those words must produce numeric actions? I study a researcher-constructed scale

model-releasesarxiv-cs-cl
22 May 2026
Hardware

Don't Collapse Your Features: Why CenterLoss Hurts OOD Detection and Multi-Scale Mahalanobis Wins

DGX agent

arXiv:2605.21493v1 Announce Type: cross Abstract: The ability to detect out-of-distribution (OOD) inputs is fundamental to safe deployment of machine learning systems. Yet, current methods often rely

hardwarearxiv-cs-cv
22 May 2026
Model Releases

Dual-Integrated Low-Latency Single-Lens Infrared Computational Imaging for Object Detection

DGX agent

arXiv:2605.21964v1 Announce Type: new Abstract: Computational imaging enables compact infrared systems, but deep-learning pipelines that combine image reconstruction and object detection often introdu

model-releasesarxiv-cs-cv
22 May 2026
Hardware

EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation

DGX agent

arXiv:2605.22051v1 Announce Type: new Abstract: Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of sp

hardwarearxiv-cs-cv
22 May 2026
Agents

Echo: Learning from Experience Data via User-Driven Refinement

DGX agent

arXiv:2605.21984v1 Announce Type: cross Abstract: Static 'human data' faces inherent limitations: it is expensive to scale and bounded by the knowledge of its creators. Continuous learning from 'exper

agentsarxiv-cs-cl
22 May 2026
Safety

Echo4DIR: 4D Implicit Heart Reconstruction from 2D Echocardiography Videos

DGX agent

arXiv:2605.22066v1 Announce Type: new Abstract: Reconstructing 4D (3D+t) cardiac geometry from sparse 2D echocardiography is highly desirable yet fundamentally challenged by geometric ambiguity and te

safetyarxiv-cs-cv
22 May 2026
Model Releases

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

DGX agent

arXiv:2605.22138v1 Announce Type: cross Abstract: How should an agent decide when and how to plan? A dominant approach builds agents as reactive policies with adaptive computation (e.g., chain-of-thou

model-releasesarxiv-cs-cl
22 May 2026
Applications

ELSA: An ELastic SNN Inference Architecture for Efficient Neuromorphic Computing

DGX agent

arXiv:2605.20802v1 Announce Type: cross Abstract: Spiking neural networks (SNNs) exploit event-driven and addition-only computation to substantially improve efficiency for intelligent computation. A k

applicationsarxiv-cs-ai
22 May 2026
Agents

Enabling Regulatory Multi-Agent Collaboration: Architecture, Challenges, and Solutions

DGX agent

arXiv:2509.09215v2 Announce Type: replace Abstract: Large language models (LLMs)-empowered autonomous agents are transforming both digital and physical environments by enabling adaptive, multi-agent c

agentsarxiv-cs-ai
22 May 2026
Safety

Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention

DGX agent

arXiv:2605.21842v1 Announce Type: cross Abstract: Standard transformer attention computes pairwise similarity between queries and keys, treating all tokens as equally salient regardless of their intri

safetyarxiv-cs-cl
22 May 2026
Research

Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning

DGX agent

arXiv:2401.00139v3 Announce Type: replace-cross Abstract: This paper introduces a causal attribution model to enhance the interpretability of large language models (LLMs) and improve their causal reas

researcharxiv-cs-cl
22 May 2026
Research

Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following

DGX agent

arXiv:2605.22607v1 Announce Type: new Abstract: Gaze following requires both scene understanding and gaze reasoning to localize the gaze target of an in-scene person. Recently, vision foundation model

researcharxiv-cs-cv
22 May 2026
Safety

Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

DGX agent

arXiv:2605.22185v1 Announce Type: new Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, thei

safetyarxiv-cs-cv
22 May 2026
Research

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding

DGX agent

arXiv:2605.22078v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video understanding tasks, yet challenges remain in efficientl

researcharxiv-cs-cv
22 May 2026
Research

EntmaxKV: Support-Aware Decoding for Entmax Attention

DGX agent

arXiv:2605.21649v1 Announce Type: cross Abstract: Long-context decoding is increasingly limited by KV-cache memory traffic since each generated token attends over a cache whose size grows linearly wit

researcharxiv-cs-cl
22 May 2026
Research

Entropy-Guided Self-Supervised Learning for Medical Image Classification

DGX agent

arXiv:2605.21970v1 Announce Type: cross Abstract: Accurate and robust medical image classification is paramount for early disease diagnosis and treatment planning. However, challenges such as limited

researcharxiv-cs-cv
22 May 2026
Research

Epicure: Navigating the Emergent Geometry of Food Ingredient Embeddings

DGX agent

arXiv:2605.22391v1 Announce Type: cross Abstract: We present Epicure, a family of three sibling skip-gram ingredient embeddings retrained from scratch on a multilingual recipe corpus. We aggregate 4.1

researcharxiv-cs-cl
22 May 2026
Model Releases

Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark

DGX agent

arXiv:2503.17599v3 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated considerable potential in general practice. However, existing benchmarks and evaluation frameworks pr

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Evaluating Commercial AI Chatbots as News Intermediaries

DGX agent

arXiv:2605.22785v1 Announce Type: new Abstract: AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their p

model-releasesarxiv-cs-cl
22 May 2026
Agents

Evaluating multimodal emotion recognition in proactive conversational agents: A user study

DGX agent

arXiv:2605.20200v1 Announce Type: cross Abstract: This article presents a multimodal emotion recognition module integrated into a proactive Socially Interactive Agent (SIA) powered by generative artif

agentsarxiv-cs-ai
22 May 2026
Model Releases

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

DGX agent

arXiv:2605.20630v1 Announce Type: new Abstract: Industrial asset operations workflows are latency-sensitive because a single user query may require coordination over sensor data, work orders, failure

model-releasesarxiv-cs-ai
22 May 2026
Research

Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents

DGX agent

arXiv:2605.22203v1 Announce Type: new Abstract: In this study, we compare the performance of four text chunking approaches: Recursive, Khmer-Aware, Sentence-Based, and LLM-Based within a Retrieval-Aug

researcharxiv-cs-cl
22 May 2026
Model Releases

Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset

DGX agent

arXiv:2605.22186v1 Announce Type: new Abstract: Event-based low-light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking t

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

EventGait: Towards Robust Gait Recognition with Event Streams

DGX agent

arXiv:2605.22139v1 Announce Type: new Abstract: Gait recognition enables non-intrusive, privacy-preserving identification but suffers in uncontrolled environments due to illumination and motion sensit

model-releasesarxiv-cs-cv
22 May 2026
Agents

EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning

DGX agent

arXiv:2605.22208v1 Announce Type: new Abstract: Multimodal Large Language Model (MLLM)-driven image restoration agent demonstrates effectiveness in degradation coupling scenarios by flexibly selecting

agentsarxiv-cs-cv
22 May 2026
Research

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control

DGX agent

arXiv:2605.21862v1 Announce Type: new Abstract: Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone. Yet r

researcharxiv-cs-ro
22 May 2026
Agents

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models

DGX agent

arXiv:2605.21931v1 Announce Type: new Abstract: Recent Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in video reasoning through reinforcement learning (RL). However, e

agentsarxiv-cs-cv
22 May 2026
Applications

Exposing Vulnerabilities in Visible-Infrared VLMs: A Unified Geometric Adversarial Framework with Cross-Task Transferability

DGX agent

arXiv:2605.22273v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, but their adversarial robustness in visible-infrared (VI

applicationsarxiv-cs-cv
22 May 2026
Research

Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention

DGX agent

arXiv:2605.22072v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for advancing complex reasoning in large language models, and

researcharxiv-cs-cl
22 May 2026
Model Releases

FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

DGX agent

arXiv:2605.22552v1 Announce Type: new Abstract: Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is

model-releasesarxiv-cs-cv
22 May 2026
Local Ai

FastTab: A Fast Table Recognizer with a Tiny Recursive Module and 1D Transformers

DGX agent

arXiv:2605.22422v1 Announce Type: new Abstract: Table structure recognition (TSR) requires both table-level coherence (row/column counts, headers, spanning cells) and precise separator localization. W

local-aiarxiv-cs-cv
22 May 2026
Model Releases

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

DGX agent

arXiv:2605.21625v1 Announce Type: cross Abstract: The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus

model-releasesarxiv-cs-cl
22 May 2026
Research

Flow-based Gaussian Splatting for Continuous-Scale Remote Sensing Image Super-Resolution

DGX agent

arXiv:2605.22147v1 Announce Type: new Abstract: High-resolution remote sensing images (RSIs) are crucial for Earth observation applications, yet acquiring them is often limited by sensor constraints a

researcharxiv-cs-cv
22 May 2026
Agents

Flying Together: Human-Guided Immersive Shared Control for Aerial Robot Teams in Unknown Environments

DGX agent

arXiv:2605.21680v1 Announce Type: new Abstract: While autonomous multi-robots can achieve safe and coordinated navigation, they often struggle to adapt to unforeseen conditions and to capture operator

agentsarxiv-cs-ro
22 May 2026
Safety

FlyRoute: Self-Evolving Agent Profiling via Data Flywheel for Adaptive Task Routing

DGX agent

arXiv:2605.22057v1 Announce Type: new Abstract: Enterprise routers assign queries to expert agents, yet deployed profiles stay static while agents evolve (prompts, tools, models), and developers rarel

safetyarxiv-cs-cl
22 May 2026
Safety

Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain

DGX agent

arXiv:2602.17186v2 Announce Type: replace Abstract: Large Vision Language Models (LVLMs) have achieved remarkable progress, yet they often suffer from language bias, producing answers without relying

safetyarxiv-cs-cv
22 May 2026
Research

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding

DGX agent

arXiv:2605.21973v1 Announce Type: new Abstract: Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream,

researcharxiv-cs-cv
22 May 2026
Research

ForeSplat: Optimization-Aware Foresight for Feed-Forward 3D Gaussian Splatting

DGX agent

arXiv:2605.22020v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) models offer fast single-pass reconstruction,but scaling them to match per-scene optimization quality is funda

researcharxiv-cs-cv
22 May 2026
Model Releases

FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments

DGX agent

arXiv:2605.22018v1 Announce Type: new Abstract: The Flooded Road Environments Dataset (FRED) is, to our knowledge, the first multi-modal autonomous driving dataset specifically targeting the collectio

model-releasesarxiv-cs-cv
22 May 2026
Safety

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

DGX agent

arXiv:2605.22671v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior

safetyarxiv-cs-cv
22 May 2026
Agents

From Automated to Autonomous: Hierarchical Agent-native Network Architecture (HANA)

DGX agent

arXiv:2605.20608v1 Announce Type: new Abstract: Realizing Level 4/5 Autonomous Networks (AN) demands a shift from static automation to agent-native intelligence. Current operations, reliant on rigid s

agentsarxiv-cs-ai
22 May 2026
← Previous
1…792793794795796…1311
Next →