AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
Human
87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
62,006 results
22 May 2026

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

Model ReleasesDGX agent

arXiv:2605.20630v1 Announce Type: new Abstract: Industrial asset operations workflows are latency-sensitive because a single user query may require coordination over sensor data, work orders, failure

Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents

ResearchDGX agent

arXiv:2605.22203v1 Announce Type: new Abstract: In this study, we compare the performance of four text chunking approaches: Recursive, Khmer-Aware, Sentence-Based, and LLM-Based within a Retrieval-Aug

Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset

Model Releases
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.22186v1 Announce Type: new Abstract: Event-based low-light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking t

EventGait: Towards Robust Gait Recognition with Event Streams

Model ReleasesDGX agent

arXiv:2605.22139v1 Announce Type: new Abstract: Gait recognition enables non-intrusive, privacy-preserving identification but suffers in uncontrolled environments due to illumination and motion sensit

EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning

AgentsDGX agent

arXiv:2605.22208v1 Announce Type: new Abstract: Multimodal Large Language Model (MLLM)-driven image restoration agent demonstrates effectiveness in degradation coupling scenarios by flexibly selecting

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control

ResearchDGX agent

arXiv:2605.21862v1 Announce Type: new Abstract: Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone. Yet r

EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models

AgentsDGX agent

arXiv:2605.21931v1 Announce Type: new Abstract: Recent Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in video reasoning through reinforcement learning (RL). However, e

Exposing Vulnerabilities in Visible-Infrared VLMs: A Unified Geometric Adversarial Framework with Cross-Task Transferability

ApplicationsDGX agent

arXiv:2605.22273v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, but their adversarial robustness in visible-infrared (VI

Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention

ResearchDGX agent

arXiv:2605.22072v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising paradigm for advancing complex reasoning in large language models, and

FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

Model ReleasesDGX agent

arXiv:2605.22552v1 Announce Type: new Abstract: Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is

FastTab: A Fast Table Recognizer with a Tiny Recursive Module and 1D Transformers

Local AiDGX agent

arXiv:2605.22422v1 Announce Type: new Abstract: Table structure recognition (TSR) requires both table-level coherence (row/column counts, headers, spanning cells) and precise separator localization. W

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

Model ReleasesDGX agent

arXiv:2605.21625v1 Announce Type: cross Abstract: The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus

Flow-based Gaussian Splatting for Continuous-Scale Remote Sensing Image Super-Resolution

ResearchDGX agent

arXiv:2605.22147v1 Announce Type: new Abstract: High-resolution remote sensing images (RSIs) are crucial for Earth observation applications, yet acquiring them is often limited by sensor constraints a

Flying Together: Human-Guided Immersive Shared Control for Aerial Robot Teams in Unknown Environments

AgentsDGX agent

arXiv:2605.21680v1 Announce Type: new Abstract: While autonomous multi-robots can achieve safe and coordinated navigation, they often struggle to adapt to unforeseen conditions and to capture operator

FlyRoute: Self-Evolving Agent Profiling via Data Flywheel for Adaptive Task Routing

SafetyDGX agent

arXiv:2605.22057v1 Announce Type: new Abstract: Enterprise routers assign queries to expert agents, yet deployed profiles stay static while agents evolve (prompts, tools, models), and developers rarel

Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain

SafetyDGX agent

arXiv:2602.17186v2 Announce Type: replace Abstract: Large Vision Language Models (LVLMs) have achieved remarkable progress, yet they often suffer from language bias, producing answers without relying

Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding

ResearchDGX agent

arXiv:2605.21973v1 Announce Type: new Abstract: Current Video-LLM approaches for Video Temporal Grounding (VTG) typically rely on direct timestamp generation from an unstructured visual-token stream,

ForeSplat: Optimization-Aware Foresight for Feed-Forward 3D Gaussian Splatting

ResearchDGX agent

arXiv:2605.22020v1 Announce Type: new Abstract: Feed-forward 3D Gaussian Splatting (3DGS) models offer fast single-pass reconstruction,but scaling them to match per-scene optimization quality is funda

FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments

Model ReleasesDGX agent

arXiv:2605.22018v1 Announce Type: new Abstract: The Flooded Road Environments Dataset (FRED) is, to our knowledge, the first multi-modal autonomous driving dataset specifically targeting the collectio

From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model

SafetyDGX agent

arXiv:2605.22671v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior

From Automated to Autonomous: Hierarchical Agent-native Network Architecture (HANA)

AgentsDGX agent

arXiv:2605.20608v1 Announce Type: new Abstract: Realizing Level 4/5 Autonomous Networks (AN) demands a shift from static automation to agent-native intelligence. Current operations, reliant on rigid s

From Baseline to Follow-Up: Counterfactual Spine DXA Image Synthesis in UK Biobank Using a Causal Hierarchical Variational Autoencoder

ResearchDGX agent

arXiv:2605.22649v1 Announce Type: new Abstract: Dual-energy X-ray absorptiometry (DXA) is widely used for large-scale skeletal assessment, yet learning controllable and interpretable factor-specific a

From Correlation to Cause: A Five-Stage Methodology for Feature Analysis in Transformer Language Models

ResearchDGX agent

arXiv:2605.22462v1 Announce Type: new Abstract: We propose a five-stage methodology for causal feature analysis in transformer language models (probe design, feature extraction, causal validation, rob

From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment

Model ReleasesDGX agent

arXiv:2605.21558v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains typically incurs high data and computational overhead. While prior efficiency efforts hav

From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning

ResearchDGX agent

arXiv:2605.22074v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards (RLVR) has shown strong promise for LLM reasoning, but outcome-based RLVR remains inefficient on hard p

From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding

Model ReleasesDGX agent

arXiv:2605.22413v1 Announce Type: new Abstract: Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation. While recent Multi

From TF-IDF to Transformers: A Comparative and Ensemble Approach to Sentiment Classification

ResearchDGX agent

arXiv:2605.22003v1 Announce Type: new Abstract: Sentiment analysis, also referred to as opinion mining, primarily tries to extract opinion from any text-based data. In the context of movie reviews and

GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation

AgentsDGX agent

arXiv:2605.22036v1 Announce Type: new Abstract: Despite significant progress in Vision-Language Navigation (VLN), existing approaches still rely on dense RGB videos that produce excessive patch tokens

GALAR-TemporalNet v2: Anatomy-Guided Dual-Branch Temporal Classification with Bidirectional Mamba and Dual-Graph GCN for Video Capsule Endoscopy -- after competition results

Local AiDGX agent

arXiv:2605.22209v1 Announce Type: new Abstract: Video Capsule Endoscopy (VCE) poses a challenging multi-label temporal classification problem, requiring simultaneous localization of 8 anatomical regio

GazePrior: Zero-Shot AR/VR Eye Tracking via Learned 3D Gaze Reconstruction

ResearchDGX agent

arXiv:2605.22359v1 Announce Type: new Abstract: Eye tracking (ET) is a foundational technology for advanced AR/VR applications. However, training ET models for every new ET device is challenging: real

General Agentic Planning Through Simulative Reasoning with World Models

AgentsDGX agent

arXiv:2507.23773v3 Announce Type: replace-cross Abstract: What does it mean to plan? Current agentic systems, whether scaffolded workflows or end-to-end policies, rely on reactive decision-making: sel

GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

SafetyDGX agent

arXiv:2605.21605v1 Announce Type: new Abstract: Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal

GenHAR: Generalizing Cross-domain Human Activity Recognition for Last-mile Delivery

ApplicationsDGX agent

arXiv:2605.22086v1 Announce Type: new Abstract: Human Activity Recognition (HAR) has shown remarkable effectiveness in various applications, such as smart healthcare and intelligent manufacturing. How

Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift

ResearchDGX agent

arXiv:2605.21849v1 Announce Type: cross Abstract: Mechanistic interpretability aims to explain a model's behavior by identifying causally responsible internal structures. Dictionary-based explainers s

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning

Model ReleasesDGX agent

arXiv:2605.22558v1 Announce Type: new Abstract: Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearan

GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations

ApplicationsDGX agent

arXiv:2605.22812v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, exi

GHI: Graphormer over Conditioned Hypergraph Incidence for Aspect-Based Sentiment Analysis

Model ReleasesDGX agent

arXiv:2605.22228v1 Announce Type: new Abstract: Aspect-based sentiment analysis (ABSA) requires models to bind sentiment evidence to the correct aspect, making it a natural testbed for fine-grained st

GLeVE: Graph-Guided Lesion Grounding with Proposal Verification in 3D CT

SafetyDGX agent

arXiv:2605.22619v1 Announce Type: new Abstract: Grounding radiology report descriptions to 3D CT volumes is essential for verifiable clinical interpretation, yet remains challenging due to the semanti

Governance by Construction for Generalist Agents

SafetyDGX agent

arXiv:2605.20874v1 Announce Type: new Abstract: Enterprise agents are increasingly expected to operate autonomously across tools and interfaces, yet production deployments require governance by constr

Governance by Design: Architecting Agentic AI for Organizational Learning and Scalable Autonomy

SafetyDGX agent

arXiv:2605.20210v1 Announce Type: cross Abstract: Agentic AI systems - systems that can pursue goals through multi-step planning and tool-mediated action with limited direct supervision - are moving f

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

Model ReleasesDGX agent

arXiv:2605.20203v1 Announce Type: cross Abstract: As older adults increasingly use LLM-based chatbots for companionship and assistance, a safety gap is emerging. Older adults may face vulnerabilities

Guided Trajectory Optimization with Sparse Scaling for Test-Time Diffusion

ResearchDGX agent

arXiv:2605.21907v1 Announce Type: new Abstract: The efficient Test-Time Scaling (TTS) paradigm offers a promising perspective for enhancing the generation performance of diffusion models. However, cur

H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning

Model ReleasesDGX agent

arXiv:2605.22629v1 Announce Type: new Abstract: Parametric human models capture global pose but cannot represent the non-rigid surface dynamics of clothing and soft tissue. Generic scene flow estimate

Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer

Model ReleasesDGX agent

arXiv:2605.22007v1 Announce Type: new Abstract: Hallucination is often viewed as a direct consequence of missing knowledge: a model answers incorrectly when the correct answer is absent from its gener

Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting

SafetyDGX agent

arXiv:2605.22258v1 Announce Type: new Abstract: Large language models (LLMs) require robust toxicity evaluation beyond explicit wording. This setting remains underexplored in Chinese, where toxicity m

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

Model ReleasesDGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

Heartbeat-Bound Hierarchical Credentials: Cryptographic Revocation for AI Agent Swarms

Local AiDGX agent

arXiv:2605.20704v1 Announce Type: cross Abstract: Autonomous AI agents that spawn sub-agent swarms create a safety gap: existing credential revocation mechanisms, OAuth~2.0 introspection, OCSP, and W3

Hierarchical Variational Policies for Reward-Guided Diffusion

SafetyDGX agent

arXiv:2605.21661v1 Announce Type: cross Abstract: Adapting pretrained diffusion models to downstream objectives such as inverse problems often requires expensive test-time guidance or optimization. We

High Quality Embeddings for Horn Logic Reasoning

ResearchDGX agent

arXiv:2605.20467v1 Announce Type: new Abstract: Neural networks can be trained to rank the choices made by logical reasoners, resulting in more efficient searches for answers. A key step in this proce

Higher Order Reasoning for Collaborative Communicationless Mobile Robot Operations

ResearchDGX agent

arXiv:2605.21901v1 Announce Type: new Abstract: In communicationless environments, multi-robot systems must operate without the constant information exchange that many coordination strategies typicall

How can reasoning capability empower the AI copilot robot in endoscopic surgery

SafetyDGX agent

arXiv:2605.22322v1 Announce Type: new Abstract: Reasoning capability has significantly advanced complex logical inference and robotic decision-making in general domains. However, its potential in the

How to Build Marcus's Algebraic Mind: Algebro-Deterministic Substrate over Galois Fields

TutorialsDGX agent

arXiv:2605.21379v2 Announce Type: cross Abstract: In The Algebraic Mind, Gary Marcus identified three components essential for any adequate cognitive architecture: operations over variables, recursive

How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing

Model ReleasesDGX agent

arXiv:2602.01851v2 Announce Type: replace Abstract: Recent generative models have achieved remarkable progress in image editing. However, existing systems and benchmarks remain largely text-guided. In

HUSKY: Humanoid Skateboarding System via Physics-Aware Whole-Body Control

TutorialsDGX agent

arXiv:2602.03205v2 Announce Type: replace Abstract: While current humanoid whole-body control frameworks predominantly rely on the static environment assumptions, addressing tasks characterized by hig

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild

Model ReleasesDGX agent

arXiv:2605.22064v1 Announce Type: new Abstract: Hy-MT2 is a family of fast-thinking multilingual translation models designed for complex real-world scenarios. It includes three model sizes: 1.8B, 7B,

HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering

Model ReleasesDGX agent

arXiv:2605.22035v1 Announce Type: cross Abstract: Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge

HyperBench: Standardizing and Scaling Synthetic Evaluation for Hyperspectral Super-Resolution

ApplicationsDGX agent

arXiv:2605.21671v1 Announce Type: cross Abstract: Hyperspectral super-resolution (HSR) reconstructs a high-spatial-resolution hyperspectral image by fusing a low-resolution hyperspectral image (LR-HSI

Hypergraph as Language

Local AiDGX agent

arXiv:2605.21858v1 Announce Type: new Abstract: Large language models (LLMs) have recently shown strong potential in modeling relational structures. However, existing approaches remain fundamentally g

IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions

Model ReleasesDGX agent

arXiv:2605.22247v1 Announce Type: new Abstract: Idioms pose a fundamental challenge for language models, as their meaning cannot be inferred from surface form alone. Understanding such expressions, th

Imagine2Real: Towards Zero-shot Humanoid-Object Interaction via Video Generative Priors

ResearchDGX agent

arXiv:2605.22272v1 Announce Type: cross Abstract: Whole-body Humanoid-Object Interaction (HOI) is bottlenecked by the scarcity of high-fidelity 3D data. While video generative priors offer a promising

← Previous
1…619620621622623…1034
Next →