AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,597 results
12 May 2026

Geometry-Aware Discretization Error of Diffusion Models

Model ReleasesDGX agent

arXiv:2605.08392v1 Announce Type: new Abstract: Practical diffusion sampling is a numerical approximation problem: under a fixed inference budget, one must simulate a reverse-time ODE or SDE using onl

Geospatial-Temporal Sensemaking of Remote Sensing Activity Detections with Multimodal Large Language Model

Model ReleasesDGX agent

arXiv:2605.10739v1 Announce Type: cross Abstract: We introduce SMART-HC-VQA, a Sentinel-2-based visual question answering dataset derived from the IARPA SMART Heavy Construction dataset, designed for

i like that there are models called bert and ernie, but in all seriousness, this update looks impressive

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

i like that there are models called bert and ernie, but in all seriousness, this update looks impressive ERNIE 5.1 is here 🚀 ERNIE 5.1 significantly reduces pretraining cost while compressing total pa

In KAME, a fast speech model starts replying instantly, while a backend LLM runs in parallel to inject deep knowledge on the fly. It’s a com…

ResearchDGX agent

In KAME, a fast speech model starts replying instantly, while a backend LLM runs in parallel to inject deep knowledge on the fly. It’s a completely different way to approach conversational AI, making

Large Language Models for Sequential Decision-Making: Improving In-Context Learning via Supervised Fine-Tuning

SafetyDGX agent

arXiv:2605.09009v1 Announce Type: cross Abstract: Large language models (LLMs) have shown remarkable in-context learning (ICL) capabilities, yet their potential for sequential decision-making remains

Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining

Model ReleasesDGX agent

arXiv:2605.10504v1 Announce Type: new Abstract: A causal-decoder block is hierarchical: lower layers build the residual basis that upper layers attend over. We identify a failure mode in GPT pretraini

LLM Agents Already Know When to Call Tools -- Even Without Reasoning

Model ReleasesDGX agent

arXiv:2605.09252v1 Announce Type: new Abstract: Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latenc

Metropolis-Adjusted Diffusion Models

SafetyDGX agent

arXiv:2605.09654v1 Announce Type: cross Abstract: Sampling from score-based diffusion models incurs bias due to both time discretisation and the approximation of the score function. A common strategy

Mismatch-Aware Adaptive Constraint Tightening for Bicycle-Model Trajectory Optimization

SafetyDGX agent

arXiv:2605.09376v1 Announce Type: new Abstract: Trajectory optimization for autonomous vehicles usually relies on the kinematic bicycle model because of its computational simplicity. However, when the

Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Models

ResearchDGX agent

arXiv:2602.01698v3 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have recently achieved strong mathematical and code reasoning performance through Reinforcement Learning (RL) post-tra

Rethinking Event-Based Object Dtection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning

ResearchDGX agent

arXiv:2605.08825v1 Announce Type: new Abstract: Event cameras provide microsecond-level temporal resolution, low latency, and high dynamic range, offering potential for perception under fast motion an

Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models

ResearchDGX agent

arXiv:2605.08145v1 Announce Type: cross Abstract: Current vision language models face hallucination and robustness issues against ambiguous or corrupted modalities. We hypothesize that these issues ca

Talked to a friend at a top AI lab. Their whole team is former journalists, training models on poems, summaries, and creative writing. I use…

IndustryDGX agent

Talked to a friend at a top AI lab. Their whole team is former journalists, training models on poems, summaries, and creative writing. I use AI every day and can see that it tends to flatten my writin

TARO: Temporal Adversarial Rectification Optimization Using Diffusion Models as Purifiers

ResearchDGX agent

arXiv:2605.08440v1 Announce Type: cross Abstract: Adversarial purification with diffusion models seeks to project adversarial examples back toward the data manifold, but balancing semantic preservatio

The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

Model ReleasesDGX agent

arXiv:2605.08427v1 Announce Type: new Abstract: Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in

Though the smartness comes with a cost: all of the prompts that were written for the old realtime voice model now need to be revised for a m…

ApplicationsDGX agent

Ethan Mollick discusses a tradeoff in OpenAI's newer realtime voice model, where improved capabilities require developers to revise prompts that were written for the previous version. The post highlig

Towards Robust Surgical Automation via Digital Twin Representations from Foundation Models

AgentsDGX agent

arXiv:2409.13107v3 Announce Type: replace Abstract: Large language model-based (LLM) agents are emerging as a powerful enabler of robust embodied intelligence due to their capability of planning compl

Training-Free Cultural Alignment of Large Language Models via Persona Disagreement

SafetyDGX agent

arXiv:2605.10843v1 Announce Type: cross Abstract: Large language models increasingly mediate decisions that turn on moral judgement, yet a growing body of evidence shows that their implicit preference

UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing

ResearchDGX agent

arXiv:2601.08321v3 Announce Type: replace Abstract: With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main

Unlocking air traffic flow prediction through microscopic aircraft-state modeling

ApplicationsDGX agent

arXiv:2605.10083v1 Announce Type: new Abstract: Short-term air traffic flow prediction in terminal airspace is essential for proactive air traffic management. Existing approaches predominantly model t

UxSID: Semantic-Aware User Interests Modeling for Ultra-Long Sequence

ResearchDGX agent

arXiv:2605.09040v1 Announce Type: new Abstract: Modeling ultra-long user sequences involves a difficult trade-off between efficiency and effectiveness. While current paradigms rely on either item-spec

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models

AgentsDGX agent

arXiv:2605.10106v1 Announce Type: cross Abstract: Recent advances in Multi-modal Large Language Models (MLLMs) target 3D spatial intelligence, yet the progress has been largely driven by post-training

ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models

ResearchDGX agent

arXiv:2510.10606v4 Announce Type: replace Abstract: Post-training Large Vision-and-Language Models (LVLMs) typically involves Supervised Fine-Tuning (SFT) for knowledge injection or Reinforcement Lear

When Child Inherits: Modeling and Exploiting Subagent Spawn in Multi-Agent Networks

AgentsDGX agent

arXiv:2605.08460v1 Announce Type: cross Abstract: Since the official release of ChatGPT in 2022, large language models (LLMs) have rapidly evolved from chatbot-style interfaces into agentic systems th

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

SafetyDGX agent

arXiv:2605.08245v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate,

When More Parameters Hurt: Foundation Model Priors Amplify Worst-Client Disparity Under Extreme Federated Heterogeneity

SafetyDGX agent

arXiv:2605.08992v1 Announce Type: new Abstract: Federated learning (FL) is increasingly used to fine-tune foundation models (FMs) on distributed private data. The community largely assumes that large-

11 May 2026

Beyond Pairs: Your Language Model is Secretly Optimizing a Preference Graph

SafetyDGX agent

arXiv:2605.08037v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) aligns language models using pairwise preference comparisons, offering a simple and effective alternative to Rein

CASCADE: Case-Based Continual Adaptation for Large Language Models During Deployment

AgentsDGX agent

arXiv:2605.06702v1 Announce Type: new Abstract: Large language models (LLMs) have become a central foundation of modern artificial intelligence, yet their lifecycle remains constrained by a rigid sepa

Code Generation and Conic Constraints for Model-Predictive Control on Microcontrollers with Conic-TinyMPC

ResearchDGX agent

arXiv:2403.18149v3 Announce Type: replace Abstract: Model-predictive control (MPC) is a state-of-the-art control method for constrained robotic systems, yet deployment on resource-limited hardware rem

Cognitive Agent Compilation for Explicit Problem Solver Modeling

SafetyDGX agent

arXiv:2605.07040v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used for tutoring, feedback generation, and content creation, but their broad pretraining makes them hard to c

Computer use with any model Hermes Agent × @trycua

AgentsDGX agent

This post from Nous Research discusses the integration of computer use capabilities with Hermes Agent models in collaboration with Claude (CUA), enabling AI agents to interact with computer interfaces

Conservative Flows: A New Paradigm of Generative Models

ResearchDGX agent

arXiv:2605.06905v1 Announce Type: new Abstract: Modern generative modeling is dominated by transport from a noise prior to data. We propose an alternative paradigm in which generation is performed by

Dataset Watermarking for Closed LLMs with Provable Detection

Model ReleasesDGX agent

arXiv:2605.06865v1 Announce Type: new Abstract: Large language models (LLMs) are pre-trained and post-trained on vast amounts of loosely curated data, raising the possibility that these models may hav

Emergence of Distortions in High-Dimensional Guided Diffusion Models

ApplicationsDGX agent

arXiv:2602.00716v4 Announce Type: replace-cross Abstract: Classifier-free guidance (CFG) is the de facto standard for conditional sampling in diffusion models, yet it often reduces sample diversity. U

Equivalence of Coarse and Fine-Grained Models for Learning with Distribution Shift

ResearchDGX agent

arXiv:2605.07005v1 Announce Type: cross Abstract: Recent work on provably efficient algorithms for learning with distribution shift has focused on two models: PQ learning (Goldwasser et al. (2020)) an

FLAM: Evaluating Model Performance with Aggregatable Measures in Federated Learning

Local AiDGX agent

arXiv:2605.07962v1 Announce Type: new Abstract: Performance evaluation is essential for assessing the quality of machine learning (ML) models and guiding deployment decisions. In federated learning (F

For browser-use AI agents, every task is dozens of model calls in a tight loop. The inference layer isn’t background infrastructure. It’s wh…

ToolsDGX agent

For browser-use AI agents, every task is dozens of model calls in a tight loop. The inference layer isn’t background infrastructure. It’s what the product runs on. @yutori_ai runs Scouts, Delegate, an

From Average Sensitivity to Small-Loss Regret Bounds under Random-Order Model

ResearchDGX agent

arXiv:2602.09457v2 Announce Type: replace-cross Abstract: We study online learning in the random-order model, where the multiset of loss functions is chosen adversarially but revealed in a uniformly r

GATO: GPU-Accelerated and Batched Trajectory Optimization for Scalable Edge Model Predictive Control

HardwareDGX agent

arXiv:2510.07625v2 Announce Type: replace Abstract: While Model Predictive Control (MPC) delivers strong performance across robotics applications, solving the underlying (batches of) nonlinear traject

HEART: Hyperspherical Embedding Alignment via Kent-Representation Traversal in Diffusion Models

SafetyDGX agent

arXiv:2605.07973v1 Announce Type: new Abstract: Text-to-image diffusion models can generate visually stunning images, yet, controlling what appears and how it appears, remains surprisingly difficult,

In modern ML accelerators, FLOPS have absolutely exploded. Often though, the bottleneck is not FLOPS but memory bandwidth. Similarly, model …

ResearchDGX agent

In modern ML accelerators, FLOPS have absolutely exploded. Often though, the bottleneck is not FLOPS but memory bandwidth. Similarly, model intelligence has exploded, causing the bottleneck to be huma

Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models

ResearchDGX agent

arXiv:2605.07721v1 Announce Type: cross Abstract: Recurrent LLM architectures have emerged as a promising approach for improving reasoning, as they enable multi-step computation in the embedding space

Mixture of Masters: Sparse Chess Language Models with Player Routing

ResearchDGX agent

arXiv:2602.04447v2 Announce Type: replace-cross Abstract: Modern chess language models are dense transformers trained on millions of games played by thousands of high-rated individuals. However, these

Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

SafetyDGX agent

arXiv:2602.07026v2 Announce Type: replace-cross Abstract: Despite the success of multimodal contrastive learning in aligning visual and linguistic representations, a persistent geometric anomaly, the

Neural CDEs as Correctors for Learned Time Series Models

ApplicationsDGX agent

arXiv:2512.12116v3 Announce Type: replace Abstract: Learned time-series models, whether continuous or discrete, are widely used for forecasting the states of dynamical systems but suffer from error ac

One Token Per Frame: Reconsidering Visual Bandwidth in World Models for VLA Policy

SafetyDGX agent

arXiv:2605.07931v1 Announce Type: cross Abstract: Vision-language-action (VLA) models increasingly rely on auxiliary world modules to plan over long horizons, yet how such modules should be parameteri

real time model streaming audio/video/text in and out with tool use

AgentsDGX agent

real time model streaming audio/video/text in and out with tool use People talk, listen, watch, think, and collaborate at the same time, in real time. We've designed an AI that works with people the s

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models

SafetyDGX agent

arXiv:2605.07800v1 Announce Type: new Abstract: Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions spe

Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models

ResearchDGX agent

arXiv:2509.25584v2 Announce Type: replace Abstract: Vision-language models achieve incredible performance across a wide range of tasks, but their large size makes inference costly. Recent work has sho

Sources: the White House's Office of the National Cyber Director and Commerce Department's CAISI are fighting over which agency should lead AI model evaluations (Washington Post)

SafetyDGX agent

Washington Post: Sources: the White House's Office of the National Cyber Director and Commerce Department's CAISI are fighting over which agency should lead AI model evaluations — As the White House g

TextLDM: Language Modeling with Continuous Latent Diffusion

SafetyDGX agent

arXiv:2605.07748v1 Announce Type: new Abstract: Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next st

The inability of AI models to produce creative variation is a huge gap. The fact that they generate similar ideas limits their ability to do…

ApplicationsDGX agent

The inability of AI models to produce creative variation is a huge gap. The fact that they generate similar ideas limits their ability to do science & the same-y writing limits their usefulness in man

The US Commerce Department removed from its website details about its May 5 agreement with Google, xAI, and Microsoft to test their AI models (Courtney Rozen/Reuters)

ApplicationsDGX agent

Courtney Rozen / Reuters: The US Commerce Department removed from its website details about its May 5 agreement with Google, xAI, and Microsoft to test their AI models — The U.S. Commerce Department r

TimeLesSeg: Unified Contrast-Agnostic Cross-Sectional and Longitudinal MS Lesion Segmentation via a Stochastic Generative Model

ResearchDGX agent

arXiv:2605.07955v1 Announce Type: cross Abstract: Multiple sclerosis (MS) expresses substantial clinical and radiological heterogeneity, which poses significant challenges for automatic lesion segment

Topology-Enhanced Alignment for Large Language Models: Trajectory Topology Loss and Topological Preference Optimization

Local AiDGX agent

arXiv:2605.07172v1 Announce Type: new Abstract: Alignment of large language models (LLMs) via SFT and RLHF/DPO typically ignores the global geometry of the representation space, relying instead on loc

Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions

ResearchDGX agent

arXiv:2605.07271v1 Announce Type: cross Abstract: Layer pruning efficiently reduces Large Language Model (LLM) computational costs but often triggers sudden performance collapse. Existing representati

Vaporizer: Breaking Watermarking Schemes for Large Language Model Outputs

ApplicationsDGX agent

arXiv:2605.07481v1 Announce Type: cross Abstract: In this paper, we investigate the recent state-of-the-art schemes for watermarking large language models (LLMs) outputs. These techniques are claimed

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing

ResearchDGX agent

arXiv:2605.06765v1 Announce Type: cross Abstract: Human speech conveys expressiveness beyond linguistic content, including personality, mood, or performance elements, such as a comforting tone or humm

9 May 2026

Frontier labs are betting AGI models will be so good you won't ever want to customize them. We think different. Building on a closed platfor…

ToolsDGX agent

Frontier labs are betting AGI models will be so good you won't ever want to customize them. We think different. Building on a closed platform means renting your intelligence. The landlord sets the ter

state management, observability, retries, permissioning, recovery paths, eval drift, human escalation. the model is only one component now.

AgentsDGX agent

This post from Harrison Chase discusses how AI agents have evolved beyond just the underlying model, emphasizing critical infrastructure components including state management, observability, retry mec

← Previous
1…187188189190191…1010
Next →