AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
31 Jul 2026

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

Model ReleasesDGX agent

arXiv:2607.27109v2 Announce Type: cross Abstract: With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-gra

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

Model ReleasesDGX agent

arXiv:2607.26465v1 Announce Type: new Abstract: Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform seq

New research from Microsoft. This one is on training computer-use agents at scale. Recent pipelines generate synthetic environments in bulk,…

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

New research from Microsoft. This one is on training computer-use agents at scale. Recent pipelines generate synthetic environments in bulk, which moved the bottleneck from how many exist to what is i

ProgFormer: Hierarchical Voxel Diffusion Transformer for Longitudinal Brain MRI Prediction

ResearchDGX agent

arXiv:2607.27537v1 Announce Type: new Abstract: Predicting future structural MRI of a brain is challenging because longitudinal changes are often subtle and confined to specific anatomical regions, wh

Recursive transformers for semiconductor thermo-mechanical reliability

Model ReleasesDGX agent

arXiv:2607.27251v1 Announce Type: new Abstract: Transformer-based surrogate models are increasingly used to replace expensive first-principles simulation in engineering design. But conventional transf

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Model ReleasesDGX agent

arXiv:2607.28509v1 Announce Type: new Abstract: Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference

ROCS: Request-Oriented Compute Sharing for Efficient Large-Scale Recommendation

HardwareDGX agent

arXiv:2607.27744v1 Announce Type: new Abstract: Modern recommendation models gain prediction quality by scaling feature-interaction and sequence modules, but production cost constraints cap how far sy

UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks

Model ReleasesDGX agent

arXiv:2607.26724v1 Announce Type: new Abstract: Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a limited set o

VESTIGE: A Knowledge-Guided Masking Strategy for Corruption-Aware Fine-Tuning of Genomic Transformers, Validated on Ancient DNA Reconstruction

Model ReleasesDGX agent

arXiv:2607.27712v1 Announce Type: new Abstract: Standard masked-language-model fine-tuning applies a uniform masking probability across every token position, assuming reconstruction difficulty is posi

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

AgentsDGX agent

arXiv:2607.27380v1 Announce Type: new Abstract: Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal ev

ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA

ApplicationsDGX agent

arXiv:2607.28442v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled new possibilities for 3D question answering (3D-QA), a ke

What Is The Performance Ceiling of My Classifier? Utilizing Category-Wise Influence Functions for Pareto Frontier Analysis

ResearchDGX agent

arXiv:2510.03950v2 Announce Type: replace Abstract: Data-centric learning seeks to improve model performance from the perspective of data quality, and has been drawing increasing attention in the mach

30 Jul 2026

AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing

Model ReleasesDGX agent

arXiv:2601.11007v2 Announce Type: replace-cross Abstract: LLM role-playing aims to portray arbitrary characters in interactive narratives, yet existing systems often suffer from limited immersion and

Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time

SafetyDGX agent

arXiv:2607.26647v1 Announce Type: new Abstract: While text-to-image diffusion models achieve impressive visual quality, they frequently struggle to maintain precise alignment with complex compositiona

BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

Model ReleasesDGX agent

arXiv:2606.19651v2 Announce Type: replace-cross Abstract: Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could augment under-represented

DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English

Model ReleasesDGX agent

arXiv:2601.22888v4 Announce Type: replace Abstract: More than 80% of the 1.6B English speakers do not use Standard American English (SAE), yet LLMs often fail to correctly identify non-SAE dialects an

Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine Translation

Model ReleasesDGX agent

arXiv:2607.26286v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as general-purpose translation systems, but their behavior is usually evaluated under a single prompt

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing

Model ReleasesDGX agent

arXiv:2607.26432v1 Announce Type: new Abstract: Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-grounded evidence f

From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs

Model ReleasesDGX agent

arXiv:2607.26571v1 Announce Type: new Abstract: The operational energy consumption of large language model (LLM) inference is becoming an increasingly important component of the environmental footprin

Hearsay: Vision-Language Medical Diagnoses Without an Image

Model ReleasesDGX agent

arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show

InferScale: GPU-Native KV Injection for Personalized LLM Serving

Model ReleasesDGX agent

arXiv:2607.27090v1 Announce Type: cross Abstract: Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histori

Investigating three real-world incidents in our cybersecurity evaluations

Model ReleasesDGX agent

Investigating three real-world incidents in our cybersecurity evaluations It happened again! This is turning into something of a pattern. Last week OpenAI accidentally exploited Hugging Face when one

Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification

Model ReleasesDGX agent

arXiv:2607.26397v1 Announce Type: new Abstract: Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: genera

Making a synthetic dataset for fine-tuning

Local AiDGX agent

I've been thinking about building a pipeline to generate reasoning training data for LLMs, but I want to avoid the common failure mode of synthetic data where you just generate the same template with

Mechanistic interpretability streamlined for everyday users like us😎 🧠

Model ReleasesDGX agent

Context: I want to give the community an Open Research (well open under Apache 2.0 clause) - tool that allows everyday users like us to look deeper into the local models we use consistently. Mechanist

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

SafetyDGX agent

arXiv:2607.27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers

ReCo: Reweighting GRPO Against Distributional Concentration

Model ReleasesDGX agent

arXiv:2607.26862v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

Model ReleasesDGX agent

arXiv:2607.26643v1 Announce Type: cross Abstract: Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applica

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Model ReleasesDGX agent

arXiv:2607.26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces

sure we lose money on every inference but we make it up in volume

Model ReleasesDGX agent

sure we lose money on every inference but we make it up in volume We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices f

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

Model ReleasesDGX agent

arXiv:2607.26355v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential

This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three sepa…

Model ReleasesDGX agent

This is absolutely wild... Anthropic reviewed their logs and found out that their own supposedly-sandboxed cyber evals had hacked three separate companies back in April without them noticing! In a rev

29 Jul 2026

From Deterministic to Generative Deep Learning for Urban Air Quality Reconstruction from Sparse Observations

Model ReleasesDGX agent

arXiv:2607.25687v1 Announce Type: cross Abstract: Full-field reconstruction of air pollution is essential for evaluating pollution exposure and supporting public health decision-making. However, the c

Generalization from Low- to Moderate-Resolution Spectra with Neural Networks for Stellar Parameter Estimation: A Case Study with DESI

Model ReleasesDGX agent

arXiv:2602.15021v2 Announce Type: replace-cross Abstract: Cross-survey generalization is a critical challenge in stellar spectral analysis, particularly in cases such as transferring from low- to mode

Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

Model ReleasesDGX agent

arXiv:2607.24898v1 Announce Type: cross Abstract: State-of-the-art toxicity detectors for text-to-image generation adopt a one-size-fits-all approach: a single universal model applying fixed safety gu

Learned, Relied Upon, or Necessary? Separating Checkpoint Dependence from Task-Level Value in Sheaf GNNs

Model ReleasesDGX agent

arXiv:2607.25387v1 Announce Type: new Abstract: Learned restriction maps in sheaf graph neural networks are often treated as proof that the model has discovered useful edge geometry. That conclusion d

MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA

Model ReleasesDGX agent

arXiv:2607.24838v1 Announce Type: cross Abstract: In medical multiple-choice question answering (MCQA), Retrieval-Augmented Generation (RAG) can supplement the domain knowledge of language models (LMs

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

Model ReleasesDGX agent

arXiv:2607.25614v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastro

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

Model ReleasesDGX agent

arXiv:2607.25915v1 Announce Type: new Abstract: Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or b

Rashomon Alignment

SafetyDGX agent

arXiv:2607.25680v1 Announce Type: cross Abstract: We propose Rashomon Alignment (RA), a new measure to assess functional similarity between two models. Existing functional similarity measures are dist

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation

Model ReleasesDGX agent

arXiv:2606.13040v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These r

Shieldstral

Model ReleasesDGX agent

arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7imes its size on text s

TabRank: Chain-of-Thought Distillation for Table Re-Rankers

Model ReleasesDGX agent

arXiv:2607.25182v1 Announce Type: cross Abstract: The ability to retrieve relevant tables for answering questions is a key task for structured information retrieval. Multi-stage retrieval systems rely

TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking

Model ReleasesDGX agent

arXiv:2607.24750v1 Announce Type: new Abstract: Large Language Models (LLMs) are temporally overexposed: trained on vast contemporary corpora, they encode present-day concepts that make them unreliabl

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

Model ReleasesDGX agent

arXiv:2607.25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), ho

28 Jul 2026

A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning

ResearchDGX agent

arXiv:2601.09624v2 Announce Type: replace-cross Abstract: Machine unlearning is becoming essential for building trustworthy and compliant language models. Yet unlearning success varies considerably ac

Adaptive Multi-Scale Forecasting and Gate-Localized Conformal Prediction for Multivariate Nonstationary Time Series

Model ReleasesDGX agent

arXiv:2607.23165v1 Announce Type: cross Abstract: We propose ABF-T-GLCP, a model-agnostic framework for forecasting and uncertainty quantification in nonstationary multivariate time series. The centra

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop

ResearchDGX agent

arXiv:2607.23002v1 Announce Type: cross Abstract: Large language models increasingly write both code and the tests meant to check it; coverage records what ran, not what was verified. We study an adve

AlloBench: Measuring Online Tool Allocation Capability in LLM Agents

Model ReleasesDGX agent

arXiv:2607.23332v1 Announce Type: new Abstract: Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

Model ReleasesDGX agent

arXiv:2607.24013v1 Announce Type: new Abstract: Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing ac

AssumptionMiner: Extracting, Tracing, and Revising Implicit Assumptions in LLM Code Generation

Model ReleasesDGX agent

arXiv:2607.22898v1 Announce Type: cross Abstract: Large language models (LLMs) generate code from natural-language prompts, yet real-world prompts rarely provide complete specifications. When prompts

Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS

Model ReleasesDGX agent

arXiv:2607.22657v1 Announce Type: cross Abstract: Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing

Compressing LLMs with MoP: Mixture of Pruners

Model ReleasesDGX agent

arXiv:2602.06127v2 Announce Type: replace Abstract: The high computational demands of Large Language Models (LLMs) motivate methods that reduce parameter count and accelerate inference. In response, m

Context-Aware Concept Distillation for Trustworthy Flood Prediction

Local AiDGX agent

arXiv:2607.23237v1 Announce Type: cross Abstract: Effective flood risk management relies on accurate forecasting, yet the 'black box' nature of stateof-the-art Deep Learning models creates a barrier t

Cost-Aware Recovery-Pathway Identification and Bayesian Optimization for Autonomous Materials Discovery

Model ReleasesDGX agent

arXiv:2607.23896v1 Announce Type: new Abstract: Autonomous laboratories automate experimental execution, but a campaign must also decide which recovery pathway merits optimization. We formulate this a

Covariance-Boosted Gaussian Processes for Spatiotemporal Irregularities

SafetyDGX agent

arXiv:2607.23018v1 Announce Type: cross Abstract: Nonstationary Gaussian process (GP) models are powerful tools for capturing input-dependent variability by adapting to observed data. However, with li

Data Pyramid for Embodied Manipulation

SafetyDGX agent

arXiv:2607.24744v1 Announce Type: cross Abstract: Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require d

Do LLMs Know Their Vulnerable Scenarios?

Model ReleasesDGX agent

arXiv:2607.23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their sa

DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification

Model ReleasesDGX agent

arXiv:2607.22644v1 Announce Type: new Abstract: Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or typ

DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

Model ReleasesDGX agent

arXiv:2607.23822v1 Announce Type: new Abstract: Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate becaus

← Previous
1…298299300301302…1036
Next →