AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,895 results
5 Jun 2026

HDST-GNN: Heterogeneous Dynamic Spatiotemporal Graph Neural Networks for Multi-Object Tracking in UAV Aerial Imagery

ResearchDGX agent

arXiv:2606.05587v1 Announce Type: new Abstract: Multi-object tracking (MOT) from UAV imagery presents unique challenges: altitude varies across sequences, objects are small and densely packed, and fre

Hierarchical Mask-Enhanced Dual Reconstruction Network for Few-Shot Fine-Grained Image Classification

ResearchDGX agent

arXiv:2506.20263v2 Announce Type: replace Abstract: Few-shot fine-grained image classification (FS-FGIC) is challenging as it requires distinguishing visually similar subclasses with extremely limited

HomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home Scenes

ResearchDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.06390v1 Announce Type: new Abstract: Indoor scene generation is crucial for robot simulation and modern interior design. However, complex layouts together with scarce 3D scene data make lea

Horse Eye Blink Detection and Classification for Equine Affective State Assessment

ResearchDGX agent

arXiv:2606.05458v1 Announce Type: new Abstract: Automated detection of equine facial action units (AUs) is a promising yet under-explored avenue for pain and affective state assessment in horses. Half

HyperVis: Continuous Latent Visual Relational Graphs on the Lorentz Hyperboloid for Compositional Reasoning

ResearchDGX agent

arXiv:2606.06100v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle with compositional reasoning that requires understanding inter-object relationships. A natural remedy is to injec

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

ResearchDGX agent

arXiv:2606.05769v1 Announce Type: new Abstract: Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize inter

InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning

ResearchDGX agent

arXiv:2603.17310v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) with extended reasoning capabilities often generate verbose and redundant reasoning traces, incurring unnecessary

InfoShield: Privacy-Preserving Speech Representations for Mental Health Screening via Information-Theoretic Optimization

ResearchDGX agent

arXiv:2606.05561v1 Announce Type: new Abstract: Speech-based mental health screening offers scalable depression detection, yet clinical deployment faces a significant barrier: users' privacy concerns

Interpreting Style Representations via Style-Eliciting Prompts

ResearchDGX agent

arXiv:2606.05716v1 Announce Type: new Abstract: Style representation learning is a powerful tool for authorship analysis and modeling writing style, yet the latent nature of learned representations ma

IR3DE: A Linear Router for Large Language Models

ResearchDGX agent

arXiv:2606.06098v1 Announce Type: new Abstract: Foundational Large Language Models (LLMs) demonstrate proficiency on a wide range of general tasks, and achieve remarkable results on various specialize

Knowledge Distillation for Visual Autoregressive Models

ResearchDGX agent

arXiv:2606.06078v1 Announce Type: new Abstract: Autoregressive (AR) image generation models are highly expressive but computationally intensive, motivating effective model compression. Knowledge disti

Latent Implicit Visual Reasoning

ResearchDGX agent

arXiv:2512.21218v2 Announce Type: replace Abstract: While Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning m

Learning Contact Representation for Leg Odometry

ResearchDGX agent

arXiv:2606.05501v1 Announce Type: new Abstract: The estimation of odometry in legged robots depends on the assumption that the velocity of the foot with respect to the world remains zero during the st

Learning from Demonstrations over Riemannian Manifolds using Neural ODEs: An Extended Abstract

ResearchDGX agent

arXiv:2606.05422v1 Announce Type: new Abstract: Learning from demonstratins (LfD) is usually performed over Euclidean spaces, while the robot state, e.g. orientation, naturally evolves over curved spa

Learning to Route LLMs from Implicit Cost-Performance Preferences via Meta-Learning

ResearchDGX agent

arXiv:2606.06178v1 Announce Type: cross Abstract: Large language models (LLMs) present a trade-off between performance and cost, where more powerful models incur greater expense. LLM routing aims to m

Learning What to Forget: Improving LLM Unlearning via Learned Token-Level Importance

ResearchDGX agent

arXiv:2606.06320v1 Announce Type: cross Abstract: Machine unlearning aims to remove targeted knowledge from a trained model while preserving its general capabilities. For autoregressive language model

LLM-Conditioned Synthesis of Pathological Gaits via Structured Gait-Language Representations

ResearchDGX agent

arXiv:2606.06048v1 Announce Type: new Abstract: Pathological gait datasets remain scarce due to privacy, recruitment, cost, and movement variability. Our work presents a multimodal LLM-guided framewor

LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

ResearchDGX agent

arXiv:2502.14145v3 Announce Type: replace Abstract: Achieving full-duplex communication in spoken dialogue systems (SDS) requires real-time coordination between listening, speaking, and thinking. This

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs

ResearchDGX agent

arXiv:2606.06286v1 Announce Type: new Abstract: Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather th

Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery

ResearchDGX agent

arXiv:2606.06267v1 Announce Type: new Abstract: Circuit discovery methods identify subgraphs that explain specific model behaviors, and structural differences between discovered circuits are commonly

MASF: A Multi-Model Adaptive Selection Framework for Abstractive Text summarization

ResearchDGX agent

arXiv:2606.05494v1 Announce Type: new Abstract: Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents a Multi-Model

Measuring the sensitivity of LLM-based structured extraction to prompt, model, and schema choices in clinical discharge summaries

ResearchDGX agent

arXiv:2606.05970v1 Announce Type: new Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream con

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering

ResearchDGX agent

arXiv:2606.05917v1 Announce Type: cross Abstract: Long-video question answering remains challenging for Vision-Language Models (VLMs), as answer-relevant evidence is often sparse, transient, and tempo

Monte Carlo Steklov Operators for Large-Scale Geometry Processing in the Wild

ResearchDGX agent

arXiv:2606.05581v1 Announce Type: cross Abstract: Intrinsic methods fill the default toolbox for geometry processing on meshes. Intrinsic operators, in particular the Laplacian, underlie methods that

MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action

ResearchDGX agent

arXiv:2606.06245v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies remain brittle in long-horizon and high-uncertainty control, where one-pass action decoding provides limited infer

MS-DKC: A Dataset Knowledge Card Framework for Designing and Adapting Medical Image Segmentation Models

ResearchDGX agent

arXiv:2606.06103v1 Announce Type: new Abstract: Medical image segmentation is often framed as a search for stronger architectures, but this can obscure a more fundamental question: what does the datas

Multi-Granularity Reasoning for Natural Language Inference

ResearchDGX agent

arXiv:2606.05181v1 Announce Type: new Abstract: Natural Language Inference (NLI) is a fundamental task in natural language understanding that requires determining the logical relationship between a pr

Multi-Task Crack Foundation Model for Engineering-Reliable Crack Representation and Topology Preservation in Civil Infrastructure

ResearchDGX agent

arXiv:2606.05641v1 Announce Type: new Abstract: Reliable crack assessment requires not only accurate pixel-level masks but also connected crack geometry and confidence estimates that remain stable und

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition

ResearchDGX agent

arXiv:2606.06065v1 Announce Type: new Abstract: Second-language (L2) speech recognition often requires transcriptions of pronunciations and intended meanings. Multi-task learning (MTL) is a natural ap

Multilingual Coreference Resolution via Cycle-Consistent Machine Translation

ResearchDGX agent

arXiv:2606.05444v1 Announce Type: new Abstract: Coreference resolution is a core NLP task, having a broad range of downstream applications, e.g.~machine translation, question answering, document summa

Multilingual Detection of Alzheimer's Disease from Speech: A Cross-Linguistic Transfer Learning Approach

ResearchDGX agent

arXiv:2606.05545v1 Announce Type: new Abstract: The development of multilingual Alzheimer's Disease Dementia (AD) detection models presents significant challenges due to the resource-intensive and tim

Multimodal Sexism Identification and Characterization using Large Language Models and Gradient Boosting

ResearchDGX agent

arXiv:2606.05997v1 Announce Type: new Abstract: We present the AILS-NTUA submission to the EXIST 2026 Lab at CLEF, addressing multimodal sexism identification and characterization in memes (Task 2) an

Narrative Knowledge Weaver: Narrative-Centric Retrieval-Augmented Reasoning for Long-Form Text Understanding

ResearchDGX agent

arXiv:2606.05724v1 Announce Type: new Abstract: Long-form narrative QA requires reasoning over evolving story worlds rather than isolated passages: answers may depend on earlier goals, changing charac

Next-Generation Parallel Decoder for LPDR: Architectural Optimization and Class-Balanced GAN-Augmentation

ResearchDGX agent

arXiv:2606.05785v1 Announce Type: new Abstract: Real-Time License Plate Detection and Recognition (LPDR) forms the backbone of modern smart cities. Although the YOLOV5-PDLPR model substantially improv

ORACLE-CT: Anatomy-Aware Support Pooling for CT Classification

ResearchDGX agent

arXiv:2606.05460v1 Announce Type: new Abstract: Abdominal CT disease classification is challenging because each scan is a large 3D volume with many possible findings, while diagnostic evidence is ofte

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding

ResearchDGX agent

arXiv:2606.06485v1 Announce Type: new Abstract: Recent advances in 3D multimodal large language models (3D-MLLMs) have enabled unified solutions for 3D scene understanding tasks, including visual ques

Parallel Jacobi Decoding for Fast Autoregressive Image Generation

ResearchDGX agent

arXiv:2606.05703v1 Announce Type: new Abstract: Autoregressive (AR) models have demonstrated remarkable performance in generating high-fidelity images. However, their inherently sequential next-token

PHUMA: Physically Reliable Humanoid Locomotion Dataset

ResearchDGX agent

arXiv:2510.26236v2 Announce Type: replace Abstract: Motion imitation is a promising approach for humanoid locomotion, enabling agents to acquire humanlike behaviors. Existing methods typically rely on

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

ResearchDGX agent

arXiv:2606.06361v1 Announce Type: new Abstract: Image-to-Video diffusion models leverage input images to generate visually stunning content, yet frequently produce motion that violates physical laws.

Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training

ResearchDGX agent

arXiv:2606.05610v1 Announce Type: new Abstract: The efficacy of continued pre-training for Large Language Models (LLMs) hinges upon hyperparameter configurations, such as learning rate and batch size.

Preserving Full 6-DOF Actuation Under Abrupt Total Rotor Failures: Passive Fault-Tolerant Flight Control Using a Biaxial-Tilt Hexacopter

ResearchDGX agent

arXiv:2606.05663v1 Announce Type: new Abstract: Conventional multirotors suffer from a rapid collapse of attainable wrench space (AWS) under abrupt total rotor failures, rendering full 6-DOF recovery

RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning

ResearchDGX agent

arXiv:2606.06033v1 Announce Type: new Abstract: Learning dexterous manipulation requires demonstrations that preserve fine hand-object interactions while remaining executable at deployment. Existing p

Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation

ResearchDGX agent

arXiv:2606.06428v1 Announce Type: new Abstract: Prior work has shown that large language models (LLMs) can translate unseen or low-resource languages by undergoing continued training or even by encodi

ReTreVal: Reasoning Tree with Validation and Cross-Problem Memory for Large Language Models

ResearchDGX agent

arXiv:2601.02880v2 Announce Type: replace-cross Abstract: Every existing inference-time reasoning framework discards all failure context at problem boundaries, leaving a model solving problem 500 no w

ReverseEOL: Improving Training-free Text Embeddings via Text Reversal in Decoder-only LLMs

ResearchDGX agent

arXiv:2606.05858v1 Announce Type: new Abstract: Recent advances in Large Language Models (LLMs) have opened new avenues for generating training-free text embeddings. However, the causal attention in d

Revising Context, Shifting Simulated Stance: Auditing LLM-Based Stance Simulation in Online Discussions

ResearchDGX agent

arXiv:2606.06443v1 Announce Type: new Abstract: Large language models are increasingly used to simulate social media users and infer how individuals may respond to online discussions. However, it rema

RQUL-UIE: Revitalizing Quality-Unstable Labels for Underwater Image Enhancement via In-Dataset Self-Supervision

ResearchDGX agent

arXiv:2606.06176v1 Announce Type: new Abstract: Underwater Image Enhancement (UIE) is essential for mitigating degradations caused by water medium. Although learning-based methods have advanced signif

SAM-Flow: Source-Anchored Masked Flow for Training-Free Image Editing

ResearchDGX agent

arXiv:2606.06228v1 Announce Type: new Abstract: Training-free image editing has recently attracted increasing attention due to its ability to modify real images using powerful pre-trained diffusion an

SC-MFJ: A Simple Haptic Quality Metric for Medical Image Segmentation

ResearchDGX agent

arXiv:2606.06199v1 Announce Type: new Abstract: Standard segmentation metrics such as Dice and Hausdorff distance measure geometric overlap but say nothing about whether a segmented surface is suitabl

Self-Learning Expression Deformations for Data-Efficient Gaussian Avatars

ResearchDGX agent

arXiv:2606.05912v1 Announce Type: new Abstract: Modeling dynamic facial expressions using 3D Gaussian representations remains challenging due to their unstructured nature. Conventional Gaussian avatar

Self-supervised Feature Disentanglement and Augmentation Network for One-class Face Anti-spoofing

ResearchDGX agent

arXiv:2503.22929v3 Announce Type: replace Abstract: Face anti-spoofing (FAS) techniques aim to enhance the security of facial identity authentication by distinguishing authentic live faces from decept

Semi-Offline Reinforcement Learning for Optimized Text Generation

ResearchDGX agent

arXiv:2306.09712v2 Announce Type: replace-cross Abstract: In reinforcement learning (RL), there are two major settings for interacting with the environment: online and offline. Online methods explore

SpanNorm: Reconciling Training Stability and Performance in Deep Transformers

ResearchDGX agent

arXiv:2601.22580v2 Announce Type: replace Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures. A critical design choice is the placeme

Tamaththul3D: High-Fidelity 3D Saudi Sign Language Avatars from Monocular Video

ResearchDGX agent

arXiv:2605.05367v2 Announce Type: replace Abstract: Existing 3D sign language avatar reconstruction methods are developed and evaluated exclusively on Western sign languages, and no 3D parametric anno

The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show

ResearchDGX agent

arXiv:2606.05328v1 Announce Type: cross Abstract: Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators. Yet

The Tell-Tale Norm: ell_2 Magnitude as a Signal for Reasoning Dynamics in Large Language Models

ResearchDGX agent

arXiv:2606.06188v1 Announce Type: new Abstract: Recent work has sought to understand Large Language Models (LLMs) reasoning, yet a principled, model-intrinsic signal that captures its layer-wise reaso

Three-Dimensional Retinal Microvasculature Restoration in OCT Angiography

ResearchDGX agent

arXiv:2606.05375v1 Announce Type: new Abstract: Optical coherence tomographic angiography (OCTA) is a powerful technique for imaging retinal microvasculature. However, acquiring reliable quantificatio

To be clear: there should not be a president of AI (let alone me 😅). Which is pretty much one point I make in the video. This reminder is a…

ResearchDGX agent

To be clear: there should not be a president of AI (let alone me 😅). Which is pretty much one point I make in the video. This reminder is about the content of the video, which is 3 years old, but ever

Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs

ResearchDGX agent

arXiv:2606.05846v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) has become a key technology for human--AI interaction. However, code-switching ASR (CS-ASR) remains particularly chal

UnHype: CLIP-Guided Hypernetworks for Dynamic LoRA Unlearning

ResearchDGX agent

arXiv:2602.03410v2 Announce Type: replace Abstract: Recent advances in large-scale diffusion models have intensified concerns about their potential misuse, particularly in generating realistic yet har

← Previous
1…177178179180181…432
Next →