AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

research

GridTimelineEvolution
19,194 results
5 Jun 2026

Comparison of Deep Learning Frameworks For Rice Disease Mapping From UAV Multispectral Imaging

ResearchDGX agent

arXiv:2606.06359v1 Announce Type: new Abstract: In this study, UAV multispectral imagery is used to segment the severity of bacterial leaf blight (BLB) in rice using convolutional neural networks (CNN

ComplexityMT: Benchmarking the Interaction Between Text Complexity and Machine Translation

ResearchDGX agent

arXiv:2606.05421v1 Announce Type: new Abstract: When a text is translated, does the translation retain the complexity of the original? We introduce ComplexityMT, a new challenge for assessing how text

Computation-Aware Event-to-Frame Reconstruction via Selective Attention

ResearchDGX agent

arXiv:2606.06142v1 Announce Type: new Abstract: Event-to-frame (E2F) reconstruction bridges asynchronous event streams with frame-based vision pipelines, but existing methods often face a trade-off be


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Congrats to Reardon and team on @flourishailabs. If they can get AI sample efficiency and energy consumption to human levels, thats going to…

ResearchDGX agent

Congrats to Reardon and team on @flourishailabs. If they can get AI sample efficiency and energy consumption to human levels, thats going to change so many things in the world! And we are live! https:

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering

ResearchDGX agent

arXiv:2510.05709v2 Announce Type: replace-cross Abstract: LLM benchmarking metrics often misstate performance and uncertainty as they rely on two assumptions that frequently do not hold in practice: (

CoT-Space: A Theoretical Framework for Internal Slow-Thinking via Reinforcement Learning

ResearchDGX agent

arXiv:2509.04027v3 Announce Type: replace-cross Abstract: Test-time scaling, primarily manifested through multi-step Chain-of-Thought (CoT) reasoning via Reinforcement Learning (RL), has emerged as a

Decomposing Factual Sycophancy in Language Models: How Size and Instruction Tuning Shape Robustness

ResearchDGX agent

arXiv:2606.06306v1 Announce Type: new Abstract: Factual sycophancy occurs when a language model abandons a correct, verifiable answer under social pressure. Because a flip occurs only when pressure to

Deep Learning-assisted AMD Staging based on OCT and OCT Angiography

ResearchDGX agent

arXiv:2606.05379v1 Announce Type: new Abstract: To develop and evaluate deep learning models for automated grading of age-related macular degeneration (AMD) severity using optical coherence tomography

Deep Learning-based 3D Oral Cavity Reconstruction Using 2D Intraoral Images

ResearchDGX agent

arXiv:2606.05998v1 Announce Type: new Abstract: Oral 3D modelling is one of the most essential stages in dentistry, and many different approaches, such as impression taking and intraoral scanning, are

DiG-Plan: Mitigating Early Commitment for Tool-Graph Planning via Diffusion Guidance

ResearchDGX agent

arXiv:2606.05728v1 Announce Type: cross Abstract: Generating executable tool plans requires selecting appropriate subsets from tool libraries, a combinatorial search problem with an exponentially larg

DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models

ResearchDGX agent

arXiv:2606.05758v1 Announce Type: new Abstract: Many modern vision-language models (VLMs) build on autoregressive decoding of discrete tokens. While text-based output interfaces enable scalable pretra

Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models

ResearchDGX agent

arXiv:2601.18383v2 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) excel at solving complex problems by explicitly generating a reasoning trace before deriving the final answer. H

Efficient Computation of Distance Functions for Navigation Vector Fields in Lie Groups

ResearchDGX agent

arXiv:2606.05372v1 Announce Type: new Abstract: Vector-field-based methods are widely used for robot control and are often applied to the path-tracking problem. Some vector field approaches require re

Emotion-Aware Image Generation from Korean Diary Text via LLM-based Prompt Translation and LoRA Fine-Tuning

ResearchDGX agent

arXiv:2606.05816v1 Announce Type: new Abstract: T2I models cannot effectively capture sentiment from various types of text, including diaries, as they primarily focus on visual object-related patterns

ExpSpeech-Net: Multimodal Fusion of Expression and Speech for Deepfake Detection

ResearchDGX agent

arXiv:2606.05760v1 Announce Type: new Abstract: Deepfake videos are increasingly challenging the credibility of online content. Many existing detection methodology relies on complex, resource-intensiv

Find an important unsolved problem you care about. Then use AI to solve it. Go deep! Talk to people. Build a community. It might take you mo…

ResearchDGX agent

Find an important unsolved problem you care about. Then use AI to solve it. Go deep! Talk to people. Build a community. It might take you months or years, but always know that AI capabilities will onl

FontFusion: Enhancing Generative Text in Diffusion Models with Typographic Conditioning

ResearchDGX agent

arXiv:2606.06066v1 Announce Type: new Abstract: Typography generation in diffusion models faces a persistent trade-off: enabling precise font control typically degrades text legibility, while maintain

FOXGLOVE: Understanding Goal-Oriented and Anchored Writing Feedback from Experts and LLMs on Argumentative Essays

ResearchDGX agent

arXiv:2606.06271v1 Announce Type: new Abstract: While large language models (LLMs) are increasingly used to generate writing feedback, there remains no systematic comparison of LLM and expert feedback

Framing, Judging, Steering: An Assessable Competency Model for Teach-ing Students to Reason With Generative AI

ResearchDGX agent

arXiv:2606.05983v1 Announce Type: cross Abstract: Generative AI makes answers easy and understanding hard, and uncritical use invites cognitive offloading. Schools still measure unaided performance, y

From Scoring to Explanations: Evaluating SHAP and LLM Rationales for Rubric-based Teaching Quality Assessment

ResearchDGX agent

arXiv:2606.05180v1 Announce Type: new Abstract: Automated scoring models are increasingly used to assign rubric-based quality ratings to complex language performances, including classroom transcripts,

Gender Artifacts from Art History to Text-to-Image Generation

ResearchDGX agent

arXiv:2606.05829v1 Announce Type: new Abstract: Artistic styles are rooted in specific socio-historical contexts that encode social hierarchies, including distinct constructions of gender. Yet in AI r

Global-Local Monte Carlo Tree Search in Vision-Language Models for Text-to-3D Indoor Scene Generation

ResearchDGX agent

arXiv:2606.06002v1 Announce Type: new Abstract: Large Vision-Language Models have achieved significant reasoning performance in various tasks.However, there are few studies on text-to-3D indoor scene

GMBFormer: An NDVI-Guided Global Memory Bank Transformer for Urban Green-Space Extraction from Ultra-High-Resolution Imagery

ResearchDGX agent

arXiv:2606.06363v1 Announce Type: new Abstract: Urban green-space extraction from ultra-high-resolution (UHR) imagery is commonly performed patch by patch, which limits semantic reuse among spatially

GRAMformer: Any-Order Modality Interactions via Volumetric Multimodal Cross-Attention

ResearchDGX agent

arXiv:2606.06249v1 Announce Type: new Abstract: Transformer-based multimodal models rely on attention mechanisms to integrate information across heterogeneous modalities. Despite their success, existi

HDST-GNN: Heterogeneous Dynamic Spatiotemporal Graph Neural Networks for Multi-Object Tracking in UAV Aerial Imagery

ResearchDGX agent

arXiv:2606.05587v1 Announce Type: new Abstract: Multi-object tracking (MOT) from UAV imagery presents unique challenges: altitude varies across sequences, objects are small and densely packed, and fre

Hierarchical Mask-Enhanced Dual Reconstruction Network for Few-Shot Fine-Grained Image Classification

ResearchDGX agent

arXiv:2506.20263v2 Announce Type: replace Abstract: Few-shot fine-grained image classification (FS-FGIC) is challenging as it requires distinguishing visually similar subclasses with extremely limited

HomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home Scenes

ResearchDGX agent

arXiv:2606.06390v1 Announce Type: new Abstract: Indoor scene generation is crucial for robot simulation and modern interior design. However, complex layouts together with scarce 3D scene data make lea

Horse Eye Blink Detection and Classification for Equine Affective State Assessment

ResearchDGX agent

arXiv:2606.05458v1 Announce Type: new Abstract: Automated detection of equine facial action units (AUs) is a promising yet under-explored avenue for pain and affective state assessment in horses. Half

HyperVis: Continuous Latent Visual Relational Graphs on the Lorentz Hyperboloid for Compositional Reasoning

ResearchDGX agent

arXiv:2606.06100v1 Announce Type: new Abstract: Vision-Language Models (VLMs) struggle with compositional reasoning that requires understanding inter-object relationships. A natural remedy is to injec

I want to offer some unsolicited advice to computer vision researchers jumping into robotics. Don't focus too much on VLMs, VLAs etc. That's…

ResearchDGX agent

I want to offer some unsolicited advice to computer vision researchers jumping into robotics. Don't focus too much on VLMs, VLAs etc. That's fine, but the real action is at the sensorimotor level. Mos

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

ResearchDGX agent

arXiv:2606.05769v1 Announce Type: new Abstract: Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize inter

InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning

ResearchDGX agent

arXiv:2603.17310v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) with extended reasoning capabilities often generate verbose and redundant reasoning traces, incurring unnecessary

InfoShield: Privacy-Preserving Speech Representations for Mental Health Screening via Information-Theoretic Optimization

ResearchDGX agent

arXiv:2606.05561v1 Announce Type: new Abstract: Speech-based mental health screening offers scalable depression detection, yet clinical deployment faces a significant barrier: users' privacy concerns

Interpreting Style Representations via Style-Eliciting Prompts

ResearchDGX agent

arXiv:2606.05716v1 Announce Type: new Abstract: Style representation learning is a powerful tool for authorship analysis and modeling writing style, yet the latent nature of learned representations ma

IR3DE: A Linear Router for Large Language Models

ResearchDGX agent

arXiv:2606.06098v1 Announce Type: new Abstract: Foundational Large Language Models (LLMs) demonstrate proficiency on a wide range of general tasks, and achieve remarkable results on various specialize

Knowledge Distillation for Visual Autoregressive Models

ResearchDGX agent

arXiv:2606.06078v1 Announce Type: new Abstract: Autoregressive (AR) image generation models are highly expressive but computationally intensive, motivating effective model compression. Knowledge disti

Latent Implicit Visual Reasoning

ResearchDGX agent

arXiv:2512.21218v2 Announce Type: replace Abstract: While Large Multimodal Models (LMMs) have made significant progress, they remain largely text-centric, relying on language as their core reasoning m

Learning Contact Representation for Leg Odometry

ResearchDGX agent

arXiv:2606.05501v1 Announce Type: new Abstract: The estimation of odometry in legged robots depends on the assumption that the velocity of the foot with respect to the world remains zero during the st

Learning from Demonstrations over Riemannian Manifolds using Neural ODEs: An Extended Abstract

ResearchDGX agent

arXiv:2606.05422v1 Announce Type: new Abstract: Learning from demonstratins (LfD) is usually performed over Euclidean spaces, while the robot state, e.g. orientation, naturally evolves over curved spa

Learning to Route LLMs from Implicit Cost-Performance Preferences via Meta-Learning

ResearchDGX agent

arXiv:2606.06178v1 Announce Type: cross Abstract: Large language models (LLMs) present a trade-off between performance and cost, where more powerful models incur greater expense. LLM routing aims to m

Learning What to Forget: Improving LLM Unlearning via Learned Token-Level Importance

ResearchDGX agent

arXiv:2606.06320v1 Announce Type: cross Abstract: Machine unlearning aims to remove targeted knowledge from a trained model while preserving its general capabilities. For autoregressive language model

Let’s go 🚀

ResearchDGX agent

Let’s go 🚀 Building AI that Builds AI: Introducing the Sakana AI RSI Lab 🚀 https://sakana.ai/rsi-lab Today, we are announcing the Sakana AI Recursive Self-Improvement (RSI) Lab: a dedicated research g

Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study

ResearchDGX agent

arXiv:2508.20693v2 Announce Type: replace-cross Abstract: Ontologies and taxonomies of research fields are critical for managing and organising scientific knowledge, as they facilitate efficient class

LLM-Conditioned Synthesis of Pathological Gaits via Structured Gait-Language Representations

ResearchDGX agent

arXiv:2606.06048v1 Announce Type: new Abstract: Pathological gait datasets remain scarce due to privacy, recruitment, cost, and movement variability. Our work presents a multimodal LLM-guided framewor

LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

ResearchDGX agent

arXiv:2502.14145v3 Announce Type: replace Abstract: Achieving full-duplex communication in spoken dialogue systems (SDS) requires real-time coordination between listening, speaking, and thinking. This

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs

ResearchDGX agent

arXiv:2606.06286v1 Announce Type: new Abstract: Large language models can reproduce training data, but existing memorization evaluations mostly measure whether models can be forced to do so, rather th

Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery

ResearchDGX agent

arXiv:2606.06267v1 Announce Type: new Abstract: Circuit discovery methods identify subgraphs that explain specific model behaviors, and structural differences between discovered circuits are commonly

MASF: A Multi-Model Adaptive Selection Framework for Abstractive Text summarization

ResearchDGX agent

arXiv:2606.05494v1 Announce Type: new Abstract: Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents a Multi-Model

Measuring the sensitivity of LLM-based structured extraction to prompt, model, and schema choices in clinical discharge summaries

ResearchDGX agent

arXiv:2606.05970v1 Announce Type: new Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream con

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering

ResearchDGX agent

arXiv:2606.05917v1 Announce Type: cross Abstract: Long-video question answering remains challenging for Vision-Language Models (VLMs), as answer-relevant evidence is often sparse, transient, and tempo

MIRAI: Prediction and Generation of High-Impact Academic Research

ResearchDGX agent

arXiv:2606.05443v1 Announce Type: cross Abstract: The rapid pace of scientific publishing has made the identification and synthesis of high-impact work an increasingly urgent challenge. We introduce M

Monte Carlo Steklov Operators for Large-Scale Geometry Processing in the Wild

ResearchDGX agent

arXiv:2606.05581v1 Announce Type: cross Abstract: Intrinsic methods fill the default toolbox for geometry processing on meshes. Intrinsic operators, in particular the Laplacian, underlie methods that

MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action

ResearchDGX agent

arXiv:2606.06245v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies remain brittle in long-horizon and high-uncertainty control, where one-pass action decoding provides limited infer

MS-DKC: A Dataset Knowledge Card Framework for Designing and Adapting Medical Image Segmentation Models

ResearchDGX agent

arXiv:2606.06103v1 Announce Type: new Abstract: Medical image segmentation is often framed as a search for stronger architectures, but this can obscure a more fundamental question: what does the datas

Multi-Granularity Reasoning for Natural Language Inference

ResearchDGX agent

arXiv:2606.05181v1 Announce Type: new Abstract: Natural Language Inference (NLI) is a fundamental task in natural language understanding that requires determining the logical relationship between a pr

Multi-Task Crack Foundation Model for Engineering-Reliable Crack Representation and Topology Preservation in Civil Infrastructure

ResearchDGX agent

arXiv:2606.05641v1 Announce Type: new Abstract: Reliable crack assessment requires not only accurate pixel-level masks but also connected crack geometry and confidence estimates that remain stable und

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition

ResearchDGX agent

arXiv:2606.06065v1 Announce Type: new Abstract: Second-language (L2) speech recognition often requires transcriptions of pronunciations and intended meanings. Multi-task learning (MTL) is a natural ap

Multilingual Coreference Resolution via Cycle-Consistent Machine Translation

ResearchDGX agent

arXiv:2606.05444v1 Announce Type: new Abstract: Coreference resolution is a core NLP task, having a broad range of downstream applications, e.g.~machine translation, question answering, document summa

Multilingual Detection of Alzheimer's Disease from Speech: A Cross-Linguistic Transfer Learning Approach

ResearchDGX agent

arXiv:2606.05545v1 Announce Type: new Abstract: The development of multilingual Alzheimer's Disease Dementia (AD) detection models presents significant challenges due to the resource-intensive and tim

Multimodal Sexism Identification and Characterization using Large Language Models and Gradient Boosting

ResearchDGX agent

arXiv:2606.05997v1 Announce Type: new Abstract: We present the AILS-NTUA submission to the EXIST 2026 Lab at CLEF, addressing multimodal sexism identification and characterization in memes (Task 2) an

← Previous
1…132133134135136…320
Next →