AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,376
  • Agents7,554
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,170
  • Local Ai4,930
  • Model Releases23,883
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,376
  • Agents7,554
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,170
  • Local Ai4,930
  • Model Releases23,883
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
Human
88,376Total entries
1Added by human
88,375Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
62,897 results
5 Jun 2026

HDST-GNN: Heterogeneous Dynamic Spatiotemporal Graph Neural Networks for Multi-Object Tracking in UAV Aerial Imagery

ResearchDGX agent

arXiv:2606.05587v1 Announce Type: new Abstract: Multi-object tracking (MOT) from UAV imagery presents unique challenges: altitude varies across sequences, objects are small and densely packed, and fre

HERO: Learning Humanoid End-Effector Control for Visual Whole-Body Open-Vocabulary Object Grasping

SafetyDGX agent

arXiv:2602.16705v3 Announce Type: replace-cross Abstract: Visual loco-manipulation of arbitrary in-the-wild objects requires accurate end-effector (EE) control and a generalizable understanding of the

Hierarchical Mask-Enhanced Dual Reconstruction Network for Few-Shot Fine-Grained Image Classification

ResearchDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2506.20263v2 Announce Type: replace Abstract: Few-shot fine-grained image classification (FS-FGIC) is challenging as it requires distinguishing visually similar subclasses with extremely limited

HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps

Model ReleasesDGX agent

arXiv:2601.02730v3 Announce Type: replace Abstract: Visual localization on standard-definition (SD) maps has emerged as a promising low-cost and scalable solution for autonomous driving. However, exis

HomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home Scenes

ResearchDGX agent

arXiv:2606.06390v1 Announce Type: new Abstract: Indoor scene generation is crucial for robot simulation and modern interior design. However, complex layouts together with scarce 3D scene data make lea

Horse Eye Blink Detection and Classification for Equine Affective State Assessment

ResearchDGX agent

arXiv:2606.05458v1 Announce Type: new Abstract: Automated detection of equine facial action units (AUs) is a promising yet under-explored avenue for pain and affective state assessment in horses. Half

Human Adults and LLMs as Scientists: Who Benefits from Active Exploration?

SafetyDGX agent

arXiv:2606.06464v1 Announce Type: new Abstract: A long-standing finding in the causal learning literature is that adults struggle to identify conjunctive causal rules, where an effect requires the sim

Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

Model ReleasesDGX agent

arXiv:2606.06388v1 Announce Type: cross Abstract: Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly pos

HyperVis: Continuous Latent Visual Relational Graphs on the Lorentz Hyperboloid for Compositional Reasoning

ResearchDGX agent

arXiv:2606.06100v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle with compositional reasoning that requires understanding inter-object relationships. A natural remedy is to injec

IA-RAG: Interval-Algebra-Driven Temporal Reasoning for Dynamic Knowledge Retrieval

Model ReleasesDGX agent

arXiv:2606.06044v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has shown strong effectiveness in grounding Large Language Models (LLMs) with external knowledge. However, existing

IDEAL: Leveraging Infinite and Dynamic Characterizations of Large Language Models for Query-focused Summarization

SafetyDGX agent

arXiv:2407.10486v3 Announce Type: replace-cross Abstract: Query-focused summarization (QFS) aims to produce summaries that answer particular questions of interest, enabling greater user control and pe

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

ResearchDGX agent

arXiv:2606.05769v1 Announce Type: new Abstract: Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize inter

Improving Answer Extraction in Context-based Question Answering Systems Using LLMs

Model ReleasesDGX agent

arXiv:2606.06197v1 Announce Type: new Abstract: Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs). However, they still face challenges in a

Improving Heart-Focused Medical Question Answering in LLMs via Variance-Aware Rubric Rewards with GRPO

Local AiDGX agent

arXiv:2606.05174v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong promise in healthcare applications. Yet deploying general-purpose models in real-world settings remains d

In-Context Multiple Instance Learning

ApplicationsDGX agent

arXiv:2606.06458v1 Announce Type: cross Abstract: Multiple Instance Learning (MIL) addresses problems where supervision is available at the level of bags of instances and has been successfully applied

InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning

ResearchDGX agent

arXiv:2603.17310v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) with extended reasoning capabilities often generate verbose and redundant reasoning traces, incurring unnecessary

InfoShield: Privacy-Preserving Speech Representations for Mental Health Screening via Information-Theoretic Optimization

ResearchDGX agent

arXiv:2606.05561v1 Announce Type: new Abstract: Speech-based mental health screening offers scalable depression detection, yet clinical deployment faces a significant barrier: users' privacy concerns

Interpreting Style Representations via Style-Eliciting Prompts

ResearchDGX agent

arXiv:2606.05716v1 Announce Type: new Abstract: Style representation learning is a powerful tool for authorship analysis and modeling writing style, yet the latent nature of learned representations ma

Inverse Design of Realizable Metasurface based Absorbers using Improved Conditioning and Diversity Enhanced Progressively Growing GANs

SafetyDGX agent

arXiv:2606.05849v1 Announce Type: cross Abstract: Metasurfaces enable precise manipulation of electromagnetic waves for applications such as beam steering, sensing, and stealth technology. However, in

Inverse Manipulation through Symbolic Planning and Residual Operator Learning

SafetyDGX agent

arXiv:2606.05248v1 Announce Type: new Abstract: Inverting a robotic task requires more than reversing symbolic state transitions or rewinding motor trajectories. In robot manipulation tasks, symbolic

IR3DE: A Linear Router for Large Language Models

ResearchDGX agent

arXiv:2606.06098v1 Announce Type: new Abstract: Foundational Large Language Models (LLMs) demonstrate proficiency on a wide range of general tasks, and achieve remarkable results on various specialize

Is Diversity All You Need for Scalable Robotic Manipulation?

SafetyDGX agent

arXiv:2507.06219v2 Announce Type: replace Abstract: Data scaling has driven remarkable success in foundation models for Natural Language Processing (NLP) and Computer Vision (CV), yet the principles o

Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing

Model ReleasesDGX agent

arXiv:2606.05172v1 Announce Type: cross Abstract: Diffusion-based image editing has achieved strong visual fidelity under natural language instructions, yet most existing systems still operate at the

Knowledge Distillation for Visual Autoregressive Models

ResearchDGX agent

arXiv:2606.06078v1 Announce Type: new Abstract: Autoregressive (AR) image generation models are highly expressive but computationally intensive, motivating effective model compression. Knowledge disti

KV-Control: Parameter-Efficient K/V Injection for Trajectory-Controlled Text-to-Motion

Model ReleasesDGX agent

arXiv:2606.05624v1 Announce Type: new Abstract: Text-conditioned 3D human motion models now synthesize plausible motions from prompts, but practical animation and embodied-agent workflows rarely stop

L-SDPPO: Policy Optimization of Spiking Diffusion Policy for Intra-vehicular Robotic Manipulation

SafetyDGX agent

arXiv:2606.06049v1 Announce Type: new Abstract: Intra-vehicular robots in spacecraft help reduce astronaut workload and improve mission efficiency. Recent research focuses on using deep learning metho

LadderMan: Learning Humanoid Perceptive Ladder Climbing

SafetyDGX agent

arXiv:2606.05873v1 Announce Type: cross Abstract: Humanoid robots hold great promise for operating in human-centered environments, yet ladder climbing remains one of the most challenging tasks due to

LANTERN: Layered Archival and Temporal Episodic Retrieval Network for Long-Context LLM Conversations

ApplicationsDGX agent

arXiv:2606.05182v1 Announce Type: new Abstract: Large language models discard critical details when conversation history is compacted to fit within finite context windows. We present LANTERN (Layered

Large Language Models are Perplexed by some Political Parties

SafetyDGX agent

arXiv:2606.05937v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used, including in political applications, but their political fairness has been little studied. We assess

Latent Implicit Visual Reasoning

ResearchDGX agent

arXiv:2512.21218v2 Announce Type: replace Abstract: While Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning m

Latent Reasoning with Normalizing Flows

SafetyDGX agent

arXiv:2606.06447v1 Announce Type: new Abstract: Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. H

LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents

Model ReleasesDGX agent

arXiv:2606.06087v1 Announce Type: new Abstract: Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substa

LeanMarathon: Toward Reliable AI Co-Mathematicians through Long-Horizon Lean Autoformalization

Local AiDGX agent

arXiv:2606.05400v1 Announce Type: cross Abstract: Long-horizon autoformalization of research mathematics fails not only at hard lemmas, but at scale: statements drift, dependencies tangle, context dec

Learning Contact Representation for Leg Odometry

ResearchDGX agent

arXiv:2606.05501v1 Announce Type: new Abstract: The estimation of odometry in legged robots depends on the assumption that the velocity of the foot with respect to the world remains zero during the st

Learning from Demonstrations over Riemannian Manifolds using Neural ODEs: An Extended Abstract

ResearchDGX agent

arXiv:2606.05422v1 Announce Type: new Abstract: Learning from demonstratins (LfD) is usually performed over Euclidean spaces, while the robot state, e.g. orientation, naturally evolves over curved spa

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models

ApplicationsDGX agent

arXiv:2606.05833v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations that fail to m

Learning of Robot Safety Policies via Adversarial Synthetic Scenarios

SafetyDGX agent

arXiv:2606.05952v1 Announce Type: new Abstract: In this work, we propose an agentic gamification framework for hazard-informed learning of robot safety policies through synthetic scenarios. We model s

Learning Predictive Visuomotor Coordination

ApplicationsDGX agent

arXiv:2503.23300v2 Announce Type: replace Abstract: Understanding and predicting human visuomotor coordination is crucial for applications in robotics, human-computer interaction, and assistive techno

Learning Self-Correction in Vision-Language Models via Rollout Augmentation

TutorialsDGX agent

arXiv:2602.08503v2 Announce Type: replace-cross Abstract: Self-correction is essential for solving complex reasoning problems in vision-language models (VLMs). However, existing reinforcement learning

Learning to Route LLMs from Implicit Cost-Performance Preferences via Meta-Learning

ResearchDGX agent

arXiv:2606.06178v1 Announce Type: cross Abstract: Large language models (LLMs) present a trade-off between performance and cost, where more powerful models incur greater expense. LLM routing aims to m

Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation

SafetyDGX agent

arXiv:2606.06076v1 Announce Type: cross Abstract: While vision-language models excel at general multimodal understanding, they still struggle with visual spatial planning. We attribute this to a perce

Learning What to Forget: Improving LLM Unlearning via Learned Token-Level Importance

ResearchDGX agent

arXiv:2606.06320v1 Announce Type: cross Abstract: Machine unlearning aims to remove targeted knowledge from a trained model while preserving its general capabilities. For autoregressive language model

Less is MoE: Trimming Experts in Domain-Specialist Language Models

Model ReleasesDGX agent

arXiv:2606.05538v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models achieve strong performance through conditional computation, but their large parameter footprint poses deployment chall

Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

SafetyDGX agent

arXiv:2606.05737v1 Announce Type: new Abstract: Diffusion-based vision-language-action (VLA) models often inherit the image-generation view: actions are generated by iterative denoising. We argue that

Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study

ResearchDGX agent

arXiv:2508.20693v2 Announce Type: replace-cross Abstract: Ontologies and taxonomies of research fields are critical for managing and organising scientific knowledge, as they facilitate efficient class

LiAuto-GeoX: Efficient Grounded Driving Transformer

Model ReleasesDGX agent

arXiv:2606.05774v1 Announce Type: new Abstract: Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for auton

LightVesselNet: An Ultra-Lightweight Sub-100K Parameter Network for Retinal Blood Vessel Segmentation

Model ReleasesDGX agent

arXiv:2606.05354v1 Announce Type: new Abstract: Retinal blood vessel segmentation plays a vital role in the early detection of diabetic retinopathy and glaucoma. While recent deep learning models have

LLM-Conditioned Synthesis of Pathological Gaits via Structured Gait-Language Representations

ResearchDGX agent

arXiv:2606.06048v1 Announce Type: new Abstract: Pathological gait datasets remain scarce due to privacy, recruitment, cost, and movement variability. Our work presents a multimodal LLM-guided framewor

LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

ResearchDGX agent

arXiv:2502.14145v3 Announce Type: replace Abstract: Achieving full-duplex communication in spoken dialogue systems (SDS) requires real-time coordination between listening, speaking, and thinking. This

LLM-Guided ANN Index Optimization for Human-Object Interaction Retrieval

Model ReleasesDGX agent

arXiv:2606.05489v1 Announce Type: new Abstract: Retrieval systems underpin modern AI applications -- spanning visual search, recommendation engines, and multi-modal question answering. Modern multi-st

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs

ResearchDGX agent

arXiv:2606.06286v1 Announce Type: new Abstract: Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather th

Localizing Prompt Ambiguity in Large Language Models with Probe-Targeted Attribution

Model ReleasesDGX agent

arXiv:2606.05486v1 Announce Type: new Abstract: Prompt ambiguity is a common source of failure in large language models, but is difficult to localize because it is a latent property of the prompt, whi

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

Model ReleasesDGX agent

arXiv:2606.05677v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon ta

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

Model ReleasesDGX agent

arXiv:2606.06042v1 Announce Type: new Abstract: Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier fie

LoRi: Low-Rank Distillation for Implicit Reasoning

Model ReleasesDGX agent

arXiv:2606.05315v1 Announce Type: new Abstract: Implicit chain-of-thought (iCoT) methods aim to internalize reasoning in large language models, but often underperform explicit CoT prompting. We empiri

Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery

ResearchDGX agent

arXiv:2606.06267v1 Announce Type: new Abstract: Circuit discovery methods identify subgraphs that explain specific model behaviors, and structural differences between discovered circuits are commonly

MARDoc: A Memory-Aware Refinement Agent Framework for Multimodal Long Document QA

AgentsDGX agent

arXiv:2606.05749v1 Announce Type: new Abstract: Iterative retrieval-reasoning agents have recently shown promise for multimodal long-document question answering. However, most existing systems maintai

MASF: A Multi-Model Adaptive Selection Framework for Abstractive Text summarization

ResearchDGX agent

arXiv:2606.05494v1 Announce Type: new Abstract: Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents a Multi-Model

MAviS: A Multimodal Conversational Assistant For Avian Species

Model ReleasesDGX agent

arXiv:2603.07294v2 Announce Type: replace Abstract: Fine-grained understanding and species-specific multimodal question answering are vital for advancing biodiversity conservation and ecological monit

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

Model ReleasesDGX agent

arXiv:2606.05177v1 Announce Type: new Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and

← Previous
1…474475476477478…1049
Next →