AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,678
  • Agents7,513
  • Applications5,367
  • Concepts5
  • Hardware1,821
  • Industry6,154
  • Local Ai4,902
  • Model Releases23,619
  • Research19,969
  • Safety13,271
  • Syntheses17
  • Tools1,674
  • Tutorials3,366

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,678
  • Agents7,513
  • Applications5,367
  • Concepts5
  • Hardware1,821
  • Industry6,154
  • Local Ai4,902
  • Model Releases23,619
  • Research19,969
  • Safety13,271
  • Syntheses17
  • Tools1,674
  • Tutorials3,366

Source
HumanDGX agent

Content type
87,678Total entries
1Added by human
87,677Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
51,512 results
Model Releases

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?

DGX agent

arXiv:2604.10718v1 Announce Type: new Abstract: Accelerating scientific discovery requires the identification of which experiments would yield the best outcomes before committing resources to costly p

model-releasesarxiv-cs-ai
14 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding

DGX agent

arXiv:2604.11244v1 Announce Type: new Abstract: Advances in Multimodal Large Language Models (MLLMs) are transforming video captioning from a descriptive endpoint into a semantic interface for both vi

model-releasesarxiv-cs-cv
14 Apr 2026
Model Releases

SpecMoE: A Fast and Efficient Mixture-of-Experts Inference via Self-Assisted Speculative Decoding

DGX agent

arXiv:2604.10152v1 Announce Type: new Abstract: The Mixture-of-Experts (MoE) architecture has emerged as a promising approach to mitigate the rising computational costs of large language models (LLMs)

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding

DGX agent

arXiv:2604.09557v1 Announce Type: cross Abstract: Speculative Decoding (SD) has emerged as a critical technique for accelerating Large Language Model (LLM) inference. Unlike deterministic system optim

model-releasesarxiv-cs-ai
14 Apr 2026
Research

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs

DGX agent

arXiv:2509.22220v2 Announce Type: replace Abstract: Prevalent semantic speech tokenizers, designed to capture linguistic content, are surprisingly fragile. We find they are not robust to meaning-irrel

researcharxiv-cs-cl
14 Apr 2026
Model Releases

STaR-DRO: Stateful Tsallis Reweighting for Group-Robust Structured Prediction

DGX agent

arXiv:2604.09737v1 Announce Type: cross Abstract: Structured prediction requires models to generate ontology-constrained labels, grounded evidence, and valid structure under ambiguity, label skew, and

model-releasesarxiv-cs-ai
14 Apr 2026
Tutorials

SVSR: A Self-Verification and Self-Rectification Paradigm for Multimodal Reasoning

DGX agent

arXiv:2604.10228v1 Announce Type: new Abstract: Current multimodal models often suffer from shallow reasoning, leading to errors caused by incomplete or inconsistent thought processes. To address this

tutorialsarxiv-cs-ai
14 Apr 2026
Research

The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise

DGX agent

arXiv:2604.09780v1 Announce Type: new Abstract: Mixture of Experts (MoEs) are now ubiquitous in large language models, yet the mechanisms behind their 'expert specialization' remain poorly understood.

researcharxiv-cs-ai
14 Apr 2026
Model Releases

Toward Generalized Cross-Lingual Hateful Language Detection with Web-Scale Data and Ensemble LLM Annotations

DGX agent

arXiv:2604.09625v1 Announce Type: new Abstract: We study whether large-scale unlabelled web data and LLM-based synthetic annotations can improve multilingual hate speech detection. Starting from texts

model-releasesarxiv-cs-cl
14 Apr 2026
Safety

Trajectory-based actuator identification via differentiable simulation

DGX agent

arXiv:2604.10351v1 Announce Type: new Abstract: Accurate actuation models are critical for bridging the gap between simulation and real robot behavior, yet obtaining high-fidelity actuator dynamics ty

safetyarxiv-cs-ro
14 Apr 2026
Model Releases

Trust Your Memory: Verifiable Control of Smart Homes through Reinforcement Learning with Multi-dimensional Rewards

DGX agent

arXiv:2604.10110v1 Announce Type: new Abstract: Large Language Models (LLMs) have become a key foundation for enabling personalized smart home experiences. While existing studies have explored how sma

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Tuning Qwen2.5-VL to Improve Its Web Interaction Skills

DGX agent

arXiv:2604.09571v1 Announce Type: cross Abstract: Recent advances in vision-language models (VLMs) have sparked growing interest in using them to automate web tasks, yet their feasibility as independe

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Turning Generators into Retrievers: Unlocking MLLMs for Natural Language-Guided Geo-Localization

DGX agent

arXiv:2604.10721v1 Announce Type: cross Abstract: Natural-language Guided Cross-view Geo-localization (NGCG) aims to retrieve geo-tagged satellite imagery using textual descriptions of ground scenes.

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Variable Selection Using Relative Importance Rankings

DGX agent

arXiv:2509.10853v2 Announce Type: replace-cross Abstract: Although conceptually related, variable selection and relative importance (RI) analysis have been treated quite differently in the literature.

model-releasesarxiv-cs-lg
14 Apr 2026
Model Releases

Virtual Smart Metering in District Heating Networks via Heterogeneous Spatial-Temporal Graph Neural Networks

DGX agent

arXiv:2604.10166v1 Announce Type: cross Abstract: Intelligent operation of thermal energy networks aims to improve energy efficiency, reliability, and operational flexibility through data-driven contr

model-releasesarxiv-cs-ai
14 Apr 2026
Research

VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites

DGX agent

arXiv:2506.14629v3 Announce Type: replace-cross Abstract: Mosquito-borne diseases pose a major global health risk, requiring early detection and proactive control of breeding sites to prevent outbreak

researcharxiv-cs-cl
14 Apr 2026
Research

Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval

DGX agent

arXiv:2604.10167v1 Announce Type: cross Abstract: Multi-vector models dominate Visual Document Retrieval (VDR) due to their fine-grained matching capabilities, but their high storage and computational

researcharxiv-cs-cl
14 Apr 2026
Safety

What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

DGX agent

arXiv:2510.26202v2 Announce Type: replace-cross Abstract: Human feedback can alter language models in unpredictable and undesirable ways, as practitioners lack a clear understanding of what feedback d

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

A Benchmark of Dexterity for Anthropomorphic Robotic Hands

DGX agent

arXiv:2604.09294v1 Announce Type: new Abstract: Dexterity is a central yet ambiguously defined concept in the design and evaluation of anthropomorphic robotic hands. In practice, the term is often use

model-releasesarxiv-cs-ro
13 Apr 2026
Research

A Closer Look at the Application of Causal Inference in Graph Representation Learning

DGX agent

arXiv:2604.08890v1 Announce Type: cross Abstract: Modeling causal relationships in graph representation learning remains a fundamental challenge. Existing approaches often draw on theories and methods

researcharxiv-cs-ai
13 Apr 2026
Model Releases

Agentic Jackal: Live Execution and Semantic Value Grounding for Text-to-JQL

DGX agent

arXiv:2604.09470v1 Announce Type: new Abstract: Translating natural language into Jira Query Language (JQL) requires resolving ambiguous field references, instance-specific categorical values, and com

model-releasesarxiv-cs-cl
13 Apr 2026
Model Releases

Anchored Sliding Window: Toward Robust and Imperceptible Linguistic Steganography

DGX agent

arXiv:2604.09066v1 Announce Type: new Abstract: Linguistic steganography based on language models typically assumes that steganographic texts are transmitted without alteration, making them fragile to

model-releasesarxiv-cs-cl
13 Apr 2026
Model Releases

Biologically-Grounded Multi-Encoder Architectures as Developability Oracles for Antibody Design

DGX agent

arXiv:2604.09369v1 Announce Type: cross Abstract: Generative models can now propose thousands of de novo antibody sequences, yet translating these designs into viable therapeutics remains constrained

model-releasesarxiv-cs-lg
13 Apr 2026
Safety

CausalVAD: De-confounding End-to-End Autonomous Driving via Causal Intervention

DGX agent

arXiv:2603.18561v2 Announce Type: replace Abstract: Planning-oriented end-to-end driving models show great promise, yet they fundamentally learn statistical correlations instead of true causal relatio

safetyarxiv-cs-cv
13 Apr 2026
Safety

Chain-of-Zoom: Extreme Super-Resolution via Scale Autoregression and Preference Alignment

DGX agent

arXiv:2505.18600v3 Announce Type: replace-cross Abstract: Modern single-image super-resolution (SISR) models deliver photo-realistic results at the scale factors on which they are trained, but collaps

safetyarxiv-cs-ai
13 Apr 2026
Tutorials

Generative View Stitching

DGX agent

arXiv:2510.24718v3 Announce Type: replace Abstract: Autoregressive video diffusion models are capable of long rollouts that are stable and consistent with history, but they are unable to guide the cur

tutorialsarxiv-cs-cv
13 Apr 2026
Model Releases

HiL-Bench (Human-in-Loop Benchmark): Do Agents Know When to Ask for Help?

DGX agent

arXiv:2604.09408v1 Announce Type: new Abstract: Frontier coding agents solve complex tasks when given complete context but collapse when specifications are incomplete or ambiguous. The bottleneck is n

model-releasesarxiv-cs-ai
13 Apr 2026
Safety

Hitem3D 2.0: Multi-View Guided Native 3D Texture Generation

DGX agent

arXiv:2604.09231v1 Announce Type: new Abstract: Although recent advances have improved the quality of 3D texture generation, existing methods still struggle with incomplete texture coverage, cross-vie

safetyarxiv-cs-cv
13 Apr 2026
Model Releases

Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation

DGX agent

arXiv:2604.08797v1 Announce Type: cross Abstract: Stories are key to transmitting values across cultures, but their interpretation varies across linguistic and cultural contexts. Thus, we introduce mu

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Mamba-Based Graph Convolutional Networks: Tackling Over-smoothing with Selective State Space

DGX agent

arXiv:2501.15461v4 Announce Type: replace Abstract: Graph Neural Networks (GNNs) have shown great success in various graph-based learning tasks. However, it often faces the issue of over-smoothing as

model-releasesarxiv-cs-lg
13 Apr 2026
Research

Music Audio-Visual Question Answering Requires Specialized Multimodal Designs

DGX agent

arXiv:2505.20638v2 Announce Type: replace-cross Abstract: While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music nec

researcharxiv-cs-cv
13 Apr 2026
Model Releases

PhysInOne: Visual Physics Learning and Reasoning in One Suite

DGX agent

arXiv:2604.09415v1 Announce Type: cross Abstract: We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike exi

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

PinpointQA: A Dataset and Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos

DGX agent

arXiv:2604.08991v1 Announce Type: cross Abstract: Small object-centric spatial understanding in indoor videos remains a significant challenge for multimodal large language models (MLLMs), despite its

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Precise Shield: Explaining and Aligning VLLM Safety via Neuron-Level Guidance

DGX agent

arXiv:2604.08881v1 Announce Type: new Abstract: In real-world deployments, Vision-Language Large Models (VLLMs) face critical challenges from multilingual and multimodal composite attacks: harmful ima

model-releasesarxiv-cs-cv
13 Apr 2026
Applications

Relational Visual Similarity

DGX agent

arXiv:2512.07833v2 Announce Type: replace-cross Abstract: Humans do not just see attribute similarity -- we also see relational similarity. An apple is like a peach because both are reddish fruit, but

applicationsarxiv-cs-ai
13 Apr 2026
Model Releases

RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation

DGX agent

arXiv:2510.17640v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable performance on complex tasks through imitation learning in recent robotic man

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

SenBen: Sensitive Scene Graphs for Explainable Content Moderation

DGX agent

arXiv:2604.08819v1 Announce Type: cross Abstract: Content moderation systems classify images as safe or unsafe but lack spatial grounding and interpretability: they cannot explain what sensitive behav

model-releasesarxiv-cs-ai
13 Apr 2026
Safety

SHIFT: Steering Hidden Intermediates in Flow Transformers

DGX agent

arXiv:2604.09213v1 Announce Type: new Abstract: Diffusion models have become leading approaches for high-fidelity image generation. Recent DiT-based diffusion models, in particular, achieve strong pro

safetyarxiv-cs-cv
13 Apr 2026
Safety

SSPO: Subsentence-level Policy Optimization

DGX agent

arXiv:2511.04256v2 Announce Type: replace Abstract: As a key component of large language model (LLM) post-training, Reinforcement Learning from Verifiable Rewards (RLVR) has substantially improved rea

safetyarxiv-cs-cl
13 Apr 2026
Safety

Think Less, Know More: State-Aware Reasoning Compression with Knowledge Guidance for Efficient Reasoning

DGX agent

arXiv:2604.09150v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex tasks by leveraging long Chain-of-Thought (CoT), but often suffer from overthinking,

safetyarxiv-cs-cl
13 Apr 2026
Model Releases

TinyNeRV: Compact Neural Video Representations via Capacity Scaling, Distillation, and Low-Precision Inference

DGX agent

arXiv:2604.09220v1 Announce Type: new Abstract: Implicit neural video representations encode entire video sequences within the parameters of a neural network and enable constant time frame reconstruct

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

Towards Lifelong Aerial Autonomy: Geometric Memory Management for Continual Visual Place Recognition in Dynamic Environments

DGX agent

arXiv:2604.09038v1 Announce Type: cross Abstract: Robust geo-localization in changing environmental conditions is critical for long-term aerial autonomy. While visual place recognition (VPR) models pe

model-releasesarxiv-cs-cv
13 Apr 2026
Model Releases

Unified Multimodal Uncertain Inference

DGX agent

arXiv:2604.08701v1 Announce Type: new Abstract: We introduce Unified Multimodal Uncertain Inference (UMUI), a multimodal inference task spanning text, audio, and video, where models must produce calib

model-releasesarxiv-cs-cv
13 Apr 2026
Agents

V-CAGE: Vision-Closed-Loop Agentic Generation Engine for Robotic Manipulation

DGX agent

arXiv:2604.09036v1 Announce Type: new Abstract: Scaling Vision-Language-Action (VLA) models requires massive datasets that are both semantically coherent and physically feasible. However, existing sce

agentsarxiv-cs-ro
13 Apr 2026
Safety

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis

DGX agent

arXiv:2604.09330v1 Announce Type: cross Abstract: Recent advances in robot foundation models trained on large-scale human teleoperation data have enabled robots to perform increasingly complex real-wo

safetyarxiv-cs-cv
13 Apr 2026
Research

VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images

DGX agent

arXiv:2604.09531v1 Announce Type: cross Abstract: Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition. One plausible contr

researcharxiv-cs-ai
13 Apr 2026
Model Releases

Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures

DGX agent

arXiv:2604.09048v1 Announce Type: cross Abstract: While the large energy consumption of Large Language Models (LLMs) is recognized by the community, system operators lack guidance for energy-efficient

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning

DGX agent

arXiv:2510.07517v5 Announce Type: replace Abstract: Multi-agent debate (MAD) aims to improve large language model (LLM) reasoning by letting multiple agents exchange answers and then aggregate their o

model-releasesarxiv-cs-ai
13 Apr 2026
← Previous
1…385386387388389…1074
Next →