AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,376
  • Agents7,554
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,170
  • Local Ai4,930
  • Model Releases23,883
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,376
  • Agents7,554
  • Applications5,409
  • Concepts5
  • Hardware1,835
  • Industry6,170
  • Local Ai4,930
  • Model Releases23,883
  • Research20,124
  • Safety13,369
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
Human
88,376Total entries
1Added by human
88,375Found by agent
12Categories

Knowledge catalogue

All entries

GridTimelineEvolution
62,897 results
5 Jun 2026

MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following

Model ReleasesDGX agent

arXiv:2606.06058v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards is ideal for multi-constraint instruction following, yet standard group-relative policy optimization (G

Measuring the sensitivity of LLM-based structured extraction to prompt, model, and schema choices in clinical discharge summaries

ResearchDGX agent

arXiv:2606.05970v1 Announce Type: new Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream con

Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads

Local AiDGX agent
DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.05843v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate remarkable proficiency on complex vision-language tasks, the mechanisms by which they extract

Membrane: A Self-Evolving Contrastive Safety Memory for LLM Agent Defense

SafetyDGX agent

arXiv:2606.05743v1 Announce Type: cross Abstract: Despite advances in safety alignment, large language models remain vulnerable to continuously evolving jailbreaks. Existing fine-tuned safety classifi

MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering

ResearchDGX agent

arXiv:2606.05917v1 Announce Type: cross Abstract: Long-video question answering remains challenging for Vision-Language Models (VLMs), as answer-relevant evidence is often sparse, transient, and tempo

Merging model-based control with multi-agent reinforcement learning for multi-agent cooperative teaming strategies

AgentsDGX agent

arXiv:2606.06011v1 Announce Type: new Abstract: In this work, we propose a framework that combines multi-agent reinforcement learning (MARL) with model-based control to achieve safe, dynamically feasi

Meridian: Metric-Semantic Primitive Matching for Cross-View Geo-Localization Beyond Urban Environments

AgentsDGX agent

arXiv:2606.06312v1 Announce Type: new Abstract: Successful robot automation requires accurate global localization to support repeatability, task planning, goal specification, and safe operation. Howev

MIRAI: Prediction and Generation of High-Impact Academic Research

ResearchDGX agent

arXiv:2606.05443v1 Announce Type: cross Abstract: The rapid pace of scientific publishing has made the identification and synthesis of high-impact work an increasingly urgent challenge. We introduce M

MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery

AgentsDGX agent

arXiv:2606.06473v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly applied to long-horizon tasks such as scientific discovery and machine learning engineering (MLE),

MoDex: A Diffusion Policy for Sequential Multi-Object Dexterous Grasping

SafetyDGX agent

arXiv:2606.05407v1 Announce Type: new Abstract: This work addresses sequentially grasping multiple objects with a single dexterous hand without releasing those already held. Most dexterous grasping me

Monte Carlo Steklov Operators for Large-Scale Geometry Processing in the Wild

ResearchDGX agent

arXiv:2606.05581v1 Announce Type: cross Abstract: Intrinsic methods fill the default toolbox for geometry processing on meshes. Intrinsic operators, in particular the Laplacian, underlie methods that

MotionDisco: Motion Discovery for Extreme Humanoid Loco-Manipulation

ApplicationsDGX agent

arXiv:2606.06139v1 Announce Type: new Abstract: We present MotionDisco, a framework that discovers contact-rich, long-horizon humanoid loco-manipulation motions from scratch, without relying on teleop

MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action

ResearchDGX agent

arXiv:2606.06245v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies remain brittle in long-horizon and high-uncertainty control, where one-pass action decoding provides limited infer

MS-DKC: A Dataset Knowledge Card Framework for Designing and Adapting Medical Image Segmentation Models

ResearchDGX agent

arXiv:2606.06103v1 Announce Type: new Abstract: Medical image segmentation is often framed as a search for stronger architectures, but this can obscure a more fundamental question: what does the datas

Multi-Granularity Reasoning for Natural Language Inference

ResearchDGX agent

arXiv:2606.05181v1 Announce Type: new Abstract: Natural Language Inference (NLI) is a fundamental task in natural language understanding that requires determining the logical relationship between a pr

Multi-Resolution Tactile Imitation Learning for Contact-Rich Robotic Manipulation

SafetyDGX agent

arXiv:2606.06281v1 Announce Type: new Abstract: Touch sensing is beneficial for solving a wide variety of manipulation tasks. While there exists a wide range of tactile sensors with different properti

Multi-Task Crack Foundation Model for Engineering-Reliable Crack Representation and Topology Preservation in Civil Infrastructure

ResearchDGX agent

arXiv:2606.05641v1 Announce Type: new Abstract: Reliable crack assessment requires not only accurate pixel-level masks but also connected crack geometry and confidence estimates that remain stable und

Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition

ResearchDGX agent

arXiv:2606.06065v1 Announce Type: new Abstract: Second-language (L2) speech recognition often requires transcriptions of pronunciations and intended meanings. Multi-task learning (MTL) is a natural ap

Multilingual Coreference Resolution via Cycle-Consistent Machine Translation

ResearchDGX agent

arXiv:2606.05444v1 Announce Type: new Abstract: Coreference resolution is a core NLP task, having a broad range of downstream applications, e.g.~machine translation, question answering, document summa

Multilingual Detection of Alzheimer's Disease from Speech: A Cross-Linguistic Transfer Learning Approach

ResearchDGX agent

arXiv:2606.05545v1 Announce Type: new Abstract: The development of multilingual Alzheimer's Disease Dementia (AD) detection models presents significant challenges due to the resource-intensive and tim

Multimodal Sexism Identification and Characterization using Large Language Models and Gradient Boosting

ResearchDGX agent

arXiv:2606.05997v1 Announce Type: new Abstract: We present the AILS-NTUA submission to the EXIST 2026 Lab at CLEF, addressing multimodal sexism identification and characterization in memes (Task 2) an

Narrative Knowledge Weaver: Narrative-Centric Retrieval-Augmented Reasoning for Long-Form Text Understanding

ResearchDGX agent

arXiv:2606.05724v1 Announce Type: new Abstract: Long-form narrative QA requires reasoning over evolving story worlds rather than isolated passages: answers may depend on earlier goals, changing charac

NAVIRA: Decoupled Stochastic Remasking for Masked Diffusion Language Models

Local AiDGX agent

arXiv:2606.06031v1 Announce Type: new Abstract: Masked diffusion language models generate text by iteratively unmasking many tokens in parallel, but this speed comes with a correction problem: tokens

Next-Generation Parallel Decoder for LPDR: Architectural Optimization and Class-Balanced GAN-Augmentation

ResearchDGX agent

arXiv:2606.05785v1 Announce Type: new Abstract: Real-Time License Plate Detection and Recognition (LPDR) forms the backbone of modern smart cities. Although the YOLOV5-PDLPR model substantially improv

NIV: Neural Axis Variations for Variable Font Generation

ResearchDGX agent

arXiv:2606.05261v1 Announce Type: new Abstract: Variable fonts enable continuous variation of glyph geometry along semantic design axes such as weight, width, slant, and optical size. However, constru

Noise-Adaptive Regularization for Robust Multi-Label Remote Sensing Image Classification

ResearchDGX agent

arXiv:2601.08446v2 Announce Type: replace Abstract: The development of reliable methods for multi-label classification (MLC) has become a prominent research direction in remote sensing (RS). As the sc

Noise-Aware Visual Representation Learning for Medical Visual Question Answering

Model ReleasesDGX agent

arXiv:2606.05535v1 Announce Type: new Abstract: Medical visual question answering (Med-VQA) has strong potential for clinical decision support by enabling AI models to interpret medical images and ans

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

AgentsDGX agent

arXiv:2602.05843v2 Announce Type: replace Abstract: The rapid advancement of Large Language Models (LLMs) has catalyzed the development of autonomous agents capable of navigating complex environments.

Oklch+: A Three-Parameter Extension of Oklab for Improved Color Difference Prediction

Model ReleasesDGX agent

arXiv:2606.05255v1 Announce Type: cross Abstract: Oklab and its cylindrical representation Oklch are widely adopted in interpolation and design workflows as perceptually motivated color spaces, but th

OLIVE: Online Low-Rank Incremental Learning for Efficient Adaptive Exoskeletons

Model ReleasesDGX agent

arXiv:2606.05234v1 Announce Type: new Abstract: Wearable exoskeleton systems hold promise for restoring mobility in individuals with physical impairments, yet most existing controllers rely on static

On Advantage Estimates for Max@K Policy Gradients

SafetyDGX agent

arXiv:2606.06080v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards is widely used for post-training reasoning models, but sparse outcome rewards make exploration difficul

OneReason Technical Report

ApplicationsDGX agent

arXiv:2606.06260v1 Announce Type: cross Abstract: Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, adve

Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection

Model ReleasesDGX agent

arXiv:2606.06481v1 Announce Type: new Abstract: As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-writt

ORACLE-CT: Anatomy-Aware Support Pooling for CT Classification

ResearchDGX agent

arXiv:2606.05460v1 Announce Type: new Abstract: Abdominal CT disease classification is challenging because each scan is a large 3D volume with many possible findings, while diagnostic evidence is ofte

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation

SafetyDGX agent

arXiv:2606.06096v1 Announce Type: cross Abstract: Policy-gradient methods usually optimize expected return, but many real world applications care about distributional properties of returns: tail risk,

Ousiometrics: The essence of meaning aligns with a power-danger-structure framework instead of valence-arousal-dominance

SafetyDGX agent

arXiv:2110.06847v3 Announce Type: replace Abstract: From work emerging through the middle of the 20th century, the essence of meaning has become widely accepted as being described by the three orthogo

Ouvia: A User-centered Framework for Measuring Usability of Speech Translation in Real-World Communication Scenarios

ApplicationsDGX agent

arXiv:2606.06177v1 Announce Type: new Abstract: Speech translation (ST) is increasingly adopted in user applications, yet its evaluation largely focuses on decontextualized testbeds and holistic quali

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding

ResearchDGX agent

arXiv:2606.06485v1 Announce Type: new Abstract: Recent advances in 3D multimodal large language models (3D-MLLMs) have enabled unified solutions for 3D scene understanding tasks, including visual ques

Parallel Jacobi Decoding for Fast Autoregressive Image Generation

ResearchDGX agent

arXiv:2606.05703v1 Announce Type: new Abstract: Autoregressive (AR) models have demonstrated remarkable performance in generating high-fidelity images. However, their inherently sequential next-token

PathWISE: Multi-Agent Cancer Pathway Triaging Ontology Learning from Clinical Flowcharts

AgentsDGX agent

arXiv:2605.25970v2 Announce Type: replace Abstract: Clinical pathways are disseminated as visual flowcharts where spatial topology, arrow direction, colour coding, and font weight encode critical tria

PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation

SafetyDGX agent

arXiv:2503.14295v3 Announce Type: replace Abstract: Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack suf

PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis

Model ReleasesDGX agent

arXiv:2606.05176v1 Announce Type: new Abstract: While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-s

Personal AI Agent for Camera Roll VQA

AgentsDGX agent

arXiv:2606.05275v1 Announce Type: new Abstract: We study the personal camera roll visual question answering setting. In this setting, a conversational AI assistant can access a user's personal camera

PHUMA: Physically Reliable Humanoid Locomotion Dataset

ResearchDGX agent

arXiv:2510.26236v2 Announce Type: replace Abstract: Motion imitation is a promising approach for humanoid locomotion, enabling agents to acquire humanlike behaviors. Existing methods typically rely on

Physics-Guided Deep Unfolding for Blind Cross-Sensor Spectral Super-Resolution via Learning the Spectral Transformation Function

Model ReleasesDGX agent

arXiv:2606.05759v1 Announce Type: new Abstract: Hyperspectral imaging provides rich spectral information for quantitative remote sensing, yet hyperspectral sensors remain costly and thus unavailable i

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

ResearchDGX agent

arXiv:2606.06361v1 Announce Type: new Abstract: Image-to-Video diffusion models leverage input images to generate visually stunning content, yet frequently produce motion that violates physical laws.

PiL-World: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation

SafetyDGX agent

arXiv:2606.05773v1 Announce Type: new Abstract: Vision-language-action (VLA) policies operate in a closed loop in real-world robot tasks: a robot observes the scene, executes an action chunk, and cond

Pitfalls of Evaluating Language Models with Open Benchmarks

SafetyDGX agent

arXiv:2507.00460v3 Announce Type: replace Abstract: Open Large Language Model (LLM) benchmarks, such as HELM and BIG-Bench, provide standardized and transparent evaluation protocols that support compa

PLAN-S: Bridging Planning with Latent Style Dynamics for Autonomous Driving World Models

AgentsDGX agent

arXiv:2606.06014v1 Announce Type: cross Abstract: Latent world models (LWMs) have strengthened end-to-end autonomous driving by forecasting compact scene dynamics for downstream planning. However, exi

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.05744v1 Announce Type: new Abstract: Spatial planning maps are central to territorial governance, translating planning objectives, regulations, and spatial strategies into visual forms for

Predict and Reconstruct: Joint Objectives for Self-Supervised Language Representation Learning

Model ReleasesDGX agent

arXiv:2606.05173v1 Announce Type: new Abstract: Masked language modelling (MLM) has been the dominant pre-training objective for text encoders since BERT, yet it encourages representations that are st

Predictable Scaling Laws of Optimal Hyperparameters for LLM Continued Pre-training

ResearchDGX agent

arXiv:2606.05610v1 Announce Type: new Abstract: The efficacy of continued pre-training for Large Language Models (LLMs) hinges upon hyperparameter configurations, such as learning rate and batch size.

Preserving Full 6-DOF Actuation Under Abrupt Total Rotor Failures: Passive Fault-Tolerant Flight Control Using a Biaxial-Tilt Hexacopter

ResearchDGX agent

arXiv:2606.05663v1 Announce Type: new Abstract: Conventional multirotors suffer from a rapid collapse of attainable wrench space (AWS) under abrupt total rotor failures, rendering full 6-DOF recovery

ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

Local AiDGX agent

arXiv:2606.06168v1 Announce Type: cross Abstract: We present ProSarc, an audio-only framework that detects sarcasm by modelling temporal prosodic incongruity, that is, the mismatch between local proso

ProSPy: A Profiling-Driven SQL-Python Agentic Framework for Enterprise Text-to-SQL

Model ReleasesDGX agent

arXiv:2606.05836v1 Announce Type: new Abstract: Large language models have substantially advanced Text-to-SQL systems, yet applying them to enterprise-scale databases remains challenging. Real-world d

QueryAgent-R1: Bridging Query Generation and Product Retrieval for E-Commerce Query Recommendation

SafetyDGX agent

arXiv:2606.05671v1 Announce Type: new Abstract: Query recommendation in e-commerce search aims to proactively suggest queries that match users' potential interests. However, existing methods mainly op

RadiusFPS: Efficient Farthest Point Sampling on CPUs and GPUs via Spherical Voxel Pruning

Local AiDGX agent

arXiv:2606.06255v1 Announce Type: cross Abstract: Point clouds are a primary sensory representation for robotic perception, underpinning LiDAR-based autonomous driving, simultaneous localization and m

RAPTOR+: A Visually Grounded Vision-Language Framework to Improve Clinical Trust and Auditability in Automated Cancer Referral Processing

Model ReleasesDGX agent

arXiv:2605.25956v2 Announce Type: replace Abstract: Urgent suspected colorectal cancer (CRC) referrals create operational bottlenecks because semi-structured clinical documents often require manual re

Real-Time Threat Detection from Surveillance Cameras using Machine Learning

SafetyDGX agent

arXiv:2606.05708v1 Announce Type: new Abstract: Ensuring public safety in densely populated urban environments remains a critical challenge, necessitating the deployment of intelligent and automated v

RealDexUMI: A Wearable Universal Manipulation Interface for Dexterous Robot Learning

ResearchDGX agent

arXiv:2606.06033v1 Announce Type: new Abstract: Learning dexterous manipulation requires demonstrations that preserve fine hand-object interactions while remaining executable at deployment. Existing p

← Previous
1…475476477478479…1049
Next →