AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “papers”

GridTimelineEvolution
12,151 results
Research

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts

DGX agent

arXiv:2603.19857v2 Announce Type: replace-cross Abstract: Recent Video-to-Audio (V2A) methods have achieved remarkable progress, enabling the synthesis of realistic, high-quality audio. However, they

researcharxiv-cs-cv
21 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

Forest Before Trees: Latent Superposition for Efficient Visual Reasoning

DGX agent

arXiv:2601.06803v2 Announce Type: replace Abstract: While Chain-of-Thought empowers Large Vision-Language Models with multi-step reasoning, explicit textual rationales suffer from an information bandw

local-aiarxiv-cs-cl
21 Apr 2026
Research

Fractal Characterization of Low-Correlation Signals in AI-Generated Image Detection

DGX agent

arXiv:2604.17268v1 Announce Type: new Abstract: AI-generated imagery has reached near-photorealistic fidelity, yet this technology poses significant threats to information security and societal trust.

researcharxiv-cs-cv
21 Apr 2026
Research

Frequency-guided Multi-level Reasoning for Scene Graph Generation in Video

DGX agent

arXiv:2604.17298v1 Announce Type: new Abstract: Video Scene Graph Generation aims to obtain structured semantic representations of objects and their relationships in videos for high-level understandin

researcharxiv-cs-cv
21 Apr 2026
Research

FSEVAL: Feature Selection Evaluation Toolbox and Dashboard

DGX agent

arXiv:2604.18227v1 Announce Type: new Abstract: Feature selection is a fundamental machine learning and data mining task, involved with discriminating redundant features from informative ones. It is a

researcharxiv-cs-lg
21 Apr 2026
Model Releases

Functional Similarity Metric for Neural Networks: Overcoming Parametric Ambiguity via Activation Region Analysis

DGX agent

arXiv:2604.16426v1 Announce Type: new Abstract: As modern deep learning architectures grow in complexity, representational ambiguity emerges as a critical barrier to their interpretability and reliabl

model-releasesarxiv-cs-lg
21 Apr 2026
Safety

Generative Semantic Communication via Alternating Dual-Domain Posterior Sampling

DGX agent

arXiv:2604.16796v1 Announce Type: new Abstract: Generative semantic communication (SemCom) harnesses pretrained generative priors to improve the perceptual quality of wireless image transmission. Exis

safetyarxiv-cs-cv
21 Apr 2026
Local Ai

Geodesic Semantic Search: Cartographic Navigation of Citation Graphs with Learned Local Riemannian Maps

DGX agent

arXiv:2602.23665v3 Announce Type: replace-cross Abstract: We present Geodesic Semantic Search (GSS), a retrieval system that learns node-specific Riemannian metrics on citation graphs to enable geomet

local-aiarxiv-cs-lg
21 Apr 2026
Research

Geometry-Guided 3D Visual Token Pruning for Video-Language Models

DGX agent

arXiv:2604.18260v1 Announce Type: new Abstract: Multimodal large language models have demonstrated remarkable capabilities in 2D vision, motivating their extension to 3D scene understanding. Recent st

researcharxiv-cs-cv
21 Apr 2026
Model Releases

GeoRC: A Benchmark for Geolocation Reasoning Chains

DGX agent

arXiv:2601.21278v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) are good at recognizing the global location of a photograph -- their geolocation prediction accuracy rivals the

model-releasesarxiv-cs-cl
21 Apr 2026
Applications

Graph Neural Networks for Graphs with Heterophily: A Survey

DGX agent

arXiv:2202.07082v4 Announce Type: replace Abstract: Recent years have witnessed fast developments of graph neural networks (GNNs) that have benefited myriad graph analytic tasks and applications. Most

applicationsarxiv-cs-lg
21 Apr 2026
Model Releases

GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling

DGX agent

arXiv:2604.18556v1 Announce Type: new Abstract: Weight quantization has become a standard tool for efficient LLM deployment, especially for local inference, where models are now routinely served at 2-

model-releasesarxiv-cs-cl
21 Apr 2026
Hardware

HF becoming the platform for agents (assisted by their humans) to use and build AI (rather than just leveraging APIs)!

DGX agent

HF becoming the platform for agents (assisted by their humans) to use and build AI (rather than just leveraging APIs)! Introducing ml-intern, the agent that just automated the post-training team @hugg

hardwareclem-delangue--x
21 Apr 2026
Safety

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

DGX agent

arXiv:2507.05920v2 Announce Type: replace Abstract: State-of-the-art large multi-modal models (LMMs) face challenges when processing high-resolution images, as these inputs are converted into enormous

safetyarxiv-cs-cv
21 Apr 2026
Research

HopWeaver: Cross-Document Synthesis of High-Quality and Authentic Multi-Hop Questions

DGX agent

arXiv:2505.15087v3 Announce Type: replace Abstract: Multi-Hop Question Answering (MHQA) is crucial for evaluating the model's capability to integrate information from diverse sources. However, creatin

researcharxiv-cs-cl
21 Apr 2026
Applications

Horizon-Aware Forecasting of Passenger Assistance Demand for Rail Station Workforce Planning

DGX agent

arXiv:2604.16464v1 Announce Type: cross Abstract: Passenger assistance services are essential for accessible rail travel, yet demand varies substantially across stations and over time, creating challe

applicationsarxiv-cs-lg
21 Apr 2026
Model Releases

How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study

DGX agent

arXiv:2505.15404v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have achieved remarkable success on reasoning-intensive tasks such as mathematics and programming. However, their enha

model-releasesarxiv-cs-cl
21 Apr 2026
Applications

Hybrid Quantum Neural Networks for Enhanced Breast Cancer Thermographic Classification: A Novel Quantum-Classical Integration Approach

DGX agent

arXiv:2604.16953v1 Announce Type: cross Abstract: Breast cancer diagnosis through thermographic image analysis remains a critical challenge in medical AI, with classical deep learning approaches facin

applicationsarxiv-cs-cv
21 Apr 2026
Applications

I have been using GPT ImageGen-2 for the past weeks I didn't think that better image-generators would be a big deal but it turns out that th…

DGX agent

I have been using GPT ImageGen-2 for the past weeks I didn't think that better image-generators would be a big deal but it turns out that there is a quality threshold I didn't expect, where you can no

applicationsethan-mollick--x
21 Apr 2026
Model Releases

IDOBE: Infectious Disease Outbreak forecasting Benchmark Ecosystem

DGX agent

arXiv:2604.18521v1 Announce Type: new Abstract: Epidemic forecasting has become an integral part of real-time infectious disease outbreak response. While collaborative ensembles composed of statistica

model-releasesarxiv-cs-lg
21 Apr 2026
Research

Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context

DGX agent

arXiv:2506.10779v2 Announce Type: replace Abstract: Classroom speech and lectures often contain named entities (NEs) such as names of people and special terminology. While automatic speech recognition

researcharxiv-cs-cl
21 Apr 2026
Research

In Search of Lost DNA Sequence Pretraining

DGX agent

arXiv:2604.16570v1 Announce Type: new Abstract: DNA sequence encoding is fundamental to gene function prediction, protein synthesis, and diverse downstream biological tasks. Despite the substantial pr

researcharxiv-cs-lg
21 Apr 2026
Research

Inference-Time Temporal Probability Smoothing for Stable Video Segmentation with SAM2 under Weak Prompts

DGX agent

arXiv:2604.17115v1 Announce Type: new Abstract: Interactive video segmentation models such as SAM2 have demonstrated strong generalization across diverse visual domains. However, under weak user super

researcharxiv-cs-cv
21 Apr 2026
Safety

Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception

DGX agent

arXiv:2604.17651v1 Announce Type: new Abstract: World models, generative AI systems that simulate how environments evolve, are transforming autonomous driving, yet all existing approaches adopt an ego

safetyarxiv-cs-cv
21 Apr 2026
Safety

Instinct vs. Reflection: Unifying Token and Verbalized Confidence in Multimodal Large Models

DGX agent

arXiv:2604.17274v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in various perception and reasoning tasks. Despite this success, ens

safetyarxiv-cs-cv
21 Apr 2026
Safety

Integrated Wheel Sensor Communication using ESP32 -- A Contribution towards a Digital Twin of the Road System

DGX agent

arXiv:2509.04061v2 Announce Type: replace Abstract: While current onboard state estimation methods are adequate for most driving and safety-related applications, they do not provide insights into the

safetyarxiv-cs-ro
21 Apr 2026
Hardware

I’ve been using ml-intern for a while, and it genuinely changed my workflow. It's super good at: - Model/Dataset discovery. - Post-Training …

DGX agent

I’ve been using ml-intern for a while, and it genuinely changed my workflow. It's super good at: - Model/Dataset discovery. - Post-Training setup iteration. - Data processing workflows. Huge shoutout

hardwareclem-delangue--x
21 Apr 2026
Model Releases

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription

DGX agent

arXiv:2502.20295v2 Announce Type: replace-cross Abstract: Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractical t

model-releasesarxiv-cs-cv
21 Apr 2026
Hardware

Karpathy's autoresearch repo started an impressive trend. Agents can now train AI models to build SoTA agentic systems. And to think this is…

DGX agent

Karpathy's autoresearch repo started an impressive trend. Agents can now train AI models to build SoTA agentic systems. And to think this is just scratching the surface. Ultimately, it boils down to g

hardwareclem-delangue--x
21 Apr 2026
Research

LBFTI: Layer-Based Facial Template Inversion for Identity-Preserving Fine-Grained Face Reconstruction

DGX agent

arXiv:2604.18358v1 Announce Type: new Abstract: In face recognition systems, facial templates are widely adopted for identity authentication due to their compliance with the data minimization principl

researcharxiv-cs-cv
21 Apr 2026
Model Releases

Let's talk parsing charts 📊📈. Last week we released ParseBench, the first document OCR benchmark for AI agents. New in ParseBench: ChartDa…

DGX agent

Let's talk parsing charts 📊📈. Last week we released ParseBench, the first document OCR benchmark for AI agents. New in ParseBench: ChartDataPointMatch. Most document look at a chart and OCR the captio

model-releasesjerry-liu--x
21 Apr 2026
Applications

Lightweight Cybersickness Detection based on User-Specific Eye and Head Tracking Data in Virtual Reality

DGX agent

arXiv:2604.17158v1 Announce Type: cross Abstract: The occurrence of cybersickness in virtual reality (VR) significantly impairs users' perception and sense of immersion. Therefore, timely detection of

applicationsarxiv-cs-lg
21 Apr 2026
Research

Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage

DGX agent

arXiv:2601.03043v3 Announce Type: replace Abstract: Large language models (LLMs) demonstrate strong capabilities across a wide range of complex tasks and are increasingly deployed at scale, placing si

researcharxiv-cs-cl
21 Apr 2026
Model Releases

LiquidTAD: An Efficient Method for Temporal Action Detection via Liquid Neural Dynamics

DGX agent

arXiv:2604.18274v1 Announce Type: new Abstract: Temporal Action Detection (TAD) in untrimmed videos is currently dominated by Transformer-based architectures. While high-performing, their quadratic co

model-releasesarxiv-cs-cv
21 Apr 2026
Agents

Live LTL Progress Tracking: Towards Task-Based Exploration

DGX agent

arXiv:2604.17106v1 Announce Type: new Abstract: Motivated by the challenge presented by non-Markovian objectives in reinforcement learning (RL), we present a novel framework to track and represent the

agentsarxiv-cs-lg
21 Apr 2026
Model Releases

LLaMA-XR: A Novel Framework for Radiology Report Generation using LLaMA and QLoRA Fine Tuning

DGX agent

arXiv:2506.03178v2 Announce Type: replace-cross Abstract: Automated radiology report generation holds significant potential to reduce radiologists' workload and enhance diagnostic accuracy. However, g

model-releasesarxiv-cs-cv
21 Apr 2026
Research

LLM-AUG: Robust Wireless Data Augmentation with In-Context Learning in Large Language Models

DGX agent

arXiv:2604.17770v1 Announce Type: new Abstract: Data scarcity remains a fundamental bottleneck in applying deep learning to wireless communication problems, particularly in scenarios where collecting

researcharxiv-cs-lg
21 Apr 2026
Model Releases

Long-CODE: Isolating Pure Long-Context as an Orthogonal Dimension in Video Evaluation

DGX agent

arXiv:2604.17428v1 Announce Type: new Abstract: As video generation models achieve unprecedented capabilities, the demand for robust video evaluation metrics becomes increasingly critical. Traditional

model-releasesarxiv-cs-cv
21 Apr 2026
Model Releases

Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training

DGX agent

arXiv:2509.11983v2 Announce Type: replace Abstract: Neural network (NN) training is inherently a large-scale matrix optimization problem, yet the matrix structure of NN parameters has long been overlo

model-releasesarxiv-cs-lg
21 Apr 2026
Research

Lower Bounds and Proximally Anchored SGD for Non-Convex Minimization Under Unbounded Variance

DGX agent

arXiv:2604.16620v1 Announce Type: new Abstract: Analysis of Stochastic Gradient Descent (SGD) and its variants typically relies on the assumption of uniformly bounded variance, a condition that freque

researcharxiv-cs-lg
21 Apr 2026
Research

LTRR: Learning To Rank Retrievers for LLMs

DGX agent

arXiv:2506.13743v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems typically rely on a single fixed retriever, despite growing evidence that no single retriever performs

researcharxiv-cs-cl
21 Apr 2026
Applications

MambaKick: Early Penalty Direction Prediction from HAR Embeddings

DGX agent

arXiv:2604.16588v1 Announce Type: new Abstract: Penalty kicks in soccer are decided under extreme time constraints, where goalkeepers benefit from anticipating shot direction from the kickers motion b

applicationsarxiv-cs-cv
21 Apr 2026
Safety

MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning

DGX agent

arXiv:2602.17550v3 Announce Type: replace Abstract: Existing Reinforcement Learning with Verifiable Rewards (RLVR) algorithms, such as GRPO, rely on rigid, uniform, and symmetric trust region mechanis

safetyarxiv-cs-lg
21 Apr 2026
Model Releases

MEDN: Motion-Emotion Feature Decoupling Network for Micro-Expression Recognition

DGX agent

arXiv:2604.17899v1 Announce Type: new Abstract: Unlike macro-expression, micro-expression does not follow a strictly consistent mapping rule between emotions and Action Units (AUs). As a result, some

model-releasesarxiv-cs-cv
21 Apr 2026
Tutorials

MetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization

DGX agent

arXiv:2602.11182v2 Announce Type: replace Abstract: Existing memory systems enable Large Language Models (LLMs) to support long-horizon human-LLM interactions by persisting historical interactions bey

tutorialsarxiv-cs-cl
21 Apr 2026
Tutorials

MM-Hand: A 21-DOF Multi-modal Modular Dexterous Robotic Hand with Remote Actuation

DGX agent

arXiv:2604.17245v1 Announce Type: new Abstract: High-DOF dexterous hands require compact actuation, rich sensing, and reliable thermal behavior, but conventional designs often occupy valuable in-hand

tutorialsarxiv-cs-ro
21 Apr 2026
Model Releases

Modeling Higher-Order Brain Interactions via a Multi-View Information Bottleneck Framework for fMRI-based Psychiatric Diagnosis

DGX agent

arXiv:2604.17713v1 Announce Type: new Abstract: Resting-state functional magnetic resonance imaging (fMRI) has emerged as a cornerstone for psychiatric diagnosis, yet most approaches rely on pairwise

model-releasesarxiv-cs-lg
21 Apr 2026
Model Releases

Modeling Multiple Support Strategies within a Single Turn for Emotional Support Conversations

DGX agent

arXiv:2604.17972v1 Announce Type: new Abstract: Emotional Support Conversation (ESC) aims to assist individuals experiencing distress by generating empathetic and supportive dialogue. While prior work

model-releasesarxiv-cs-cl
21 Apr 2026
← Previous
1…230231232233234…254
Next →