AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,515 results
6 Jun 2026

CausalPOI: Spatio-Temporal Graph-Based Causal Modeling for Cold-Start POI Check-in Forecasting

ApplicationsDGX agent

arXiv:2606.05413v1 Announce Type: cross Abstract: As urban environments continue to evolve rapidly, accurately modeling the dynamic behaviour of Points of Interest is essential for supporting data-dri

EEGDancer: Dynamic Emotion Latent Space Masked Modeling with Reinforcement Learning for EEG Continuous Emotion Prediction

TutorialsDGX agent

arXiv:2606.05855v1 Announce Type: cross Abstract: Continuous electroencephalography (EEG) emotion prediction aims to model the temporal evolution of human emotional states from EEG signals. Unlike con

In-Training Defenses against Emergent Misalignment in Language Models

TutorialsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2508.06249v3 Announce Type: replace-cross Abstract: Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

ResearchDGX agent

arXiv:2606.05378v1 Announce Type: cross Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation agai

RAG Security and Privacy: Formalizing the Threat Model and Attack Surface

ApplicationsDGX agent

arXiv:2509.20324v2 Announce Type: replace-cross Abstract: Retrieval-Augmented Generation (RAG) is an emerging approach in natural language processing that combines large language models (LLMs) with ex

RAINO: Anchoring Agents in Reality, A Systematic Review and Conceptual Framework for Realism in Agent-Based Modelling

AgentsDGX agent

arXiv:2606.05167v1 Announce Type: cross Abstract: Realism is a central yet seemingly under-theorized concept in Agent-Based Modelling. This paper presents a Systematic Literature Review, aiming to ide

5 Jun 2026

A research team that includes Huawei says it successfully used Huawei's Ascend 910C chips for DeepSeek V4 Pro model's post-training, amid increased US sanctions (Coco Feng/South China Morning Post)

Model ReleasesDGX agent

Coco Feng / South China Morning Post: A research team that includes Huawei says it successfully used Huawei's Ascend 910C chips for DeepSeek V4 Pro model's post-training, amid increased US sanctions —

AdaPLD: Adaptive Retrieval and Reuse for Efficient Model-Free Speculative Decoding

ResearchDGX agent

arXiv:2606.05742v1 Announce Type: new Abstract: Speculative decoding accelerates generation by verifying multiple drafted tokens in a single target-model forward pass, reducing sequential decoding ite

Also, a lot depends on Chinese labs continuing to ship open weights models. If they stop, the frontier falls further and further behind to t…

ApplicationsDGX agent

Also, a lot depends on Chinese labs continuing to ship open weights models. If they stop, the frontier falls further and further behind to those who want to use local/fine-tuned models. I think this i

Channel-Wise Mixed-Precision Quantization for Large Language Models

Model ReleasesDGX agent

arXiv:2410.13056v4 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated remarkable success across a wide range of language tasks, but their deployment on edge devices remain

Localizing Prompt Ambiguity in Large Language Models with Probe-Targeted Attribution

Model ReleasesDGX agent

arXiv:2606.05486v1 Announce Type: new Abstract: Prompt ambiguity is a common source of failure in large language models, but is difficult to localize because it is a latent property of the prompt, whi

Sources say xAI used Claude models for distillation and training, including using personal accounts and the intermediary service Blackbox AI after being cut off (Grace Kay/The Information)

Model ReleasesDGX agent

Grace Kay / The Information: Sources say xAI used Claude models for distillation and training, including using personal accounts and the intermediary service Blackbox AI after being cut off — SpaceX's

The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show

ResearchDGX agent

arXiv:2606.05328v1 Announce Type: cross Abstract: Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators. Yet

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation

SafetyDGX agent

arXiv:2510.23497v3 Announce Type: replace Abstract: Training vision-language models (VLMs) for complex reasoning remains a challenging task, i.a. due to the scarcity of high-quality image-text reasoni

4 Jun 2026

AdaKoop: Efficient Modeling of Nonlinear Dynamics from Nonstationary Data Streams with Koopman Operator Regression

Model ReleasesDGX agent

arXiv:2606.04930v1 Announce Type: cross Abstract: Real-time data analysis requires the ability to accurately and adaptively address nonlinear dynamics in a nonstationary data stream while preserving c

Automatic Generation of Titles for Research Papers Using Language Models

Model ReleasesDGX agent

arXiv:2606.05085v1 Announce Type: cross Abstract: The title of a research paper conveys its primary idea and, occasionally, its conclusions in a clear and concise manner. Choosing an appropriate title

BreastGPT: A Multimodal Large Language Model for the Full Spectrum of Breast Cancer Clinical Routine

Model ReleasesDGX agent

arXiv:2606.04911v1 Announce Type: cross Abstract: Breast cancer remains a leading cause of cancer-related mortality among women. Its clinical management requires multimodal reasoning across a clinical

CLAW: Learning Continuous Latent Action World Models via Adversarial Latent Regularization

ResearchDGX agent

arXiv:2606.04130v1 Announce Type: new Abstract: We introduce CLAW, a fully end-to-end self-supervised framework for learning a world model jointly with continuous latent action representations directl

Evaluating Zero-Shot and One-Shot Adaptation of Small Language Models in Leader-Follower Interaction

Model ReleasesDGX agent

arXiv:2602.23312v3 Announce Type: replace-cross Abstract: Leader-follower interaction is an important paradigm in human-robot interaction (HRI). Yet, assigning roles in real time remains challenging f

EvoPrompt: Guided Prompt Evolution for Vision-Language Models Adaptation

Model ReleasesDGX agent

arXiv:2603.09493v2 Announce Type: replace-cross Abstract: The adaptation of large-scale vision-language models (VLMs) to downstream tasks with limited labeled data remains a significant challenge. Whi

Geometry-Preserving Encoder/Decoder in Latent Generative Models

TutorialsDGX agent

arXiv:2501.09876v3 Announce Type: replace-cross Abstract: Generative modeling aims to generate new data samples that resemble a given dataset. When using diffusion models for this task, one of the mai

Introducing NVIDIA Nemotron 3 Ultra. A frontier smart open model built for long-running agents that need to plan, reason, use tools and keep…

Model ReleasesDGX agent

Introducing NVIDIA Nemotron 3 Ultra. A frontier smart open model built for long-running agents that need to plan, reason, use tools and keep working across complex coding, research and enterprise work

Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-l…

Model ReleasesDGX agent

Introducing two NVIDIA Nemotron models on Together AI: Nemotron 3 Ultra for high-throughput agentic workloads and Nemotron 3.5 ASR for low-latency multilingual speech recognition. AI natives can now b

Learning Control-Affine Reduced-Order Models via Autoencoders

ResearchDGX agent

arXiv:2606.05045v1 Announce Type: cross Abstract: We present in this paper a framework for the identification of control-affine reduced-order models (ROMs). The proposed method utilizes autoencoders (

Learning Long Range Spatio-Temporal Representations over Continuous Time Dynamic Graphs with State Space Models

Model ReleasesDGX agent

arXiv:2606.04672v1 Announce Type: cross Abstract: Continuous-time dynamic graphs (CTDGs) provide a richer framework to capture fine-grained temporal patterns in evolving relational data. Long-range in

Nemotron 3 Ultra (550B-A55B) is here - our strongest open-weight model and full training recipe to date. Heavy emphasis on real-world infere…

Model ReleasesDGX agent

Nemotron 3 Ultra (550B-A55B) is here - our strongest open-weight model and full training recipe to date. Heavy emphasis on real-world inference efficiency for long-context agentic workloads. Everythin

New Benchmarking Shows Limited Generalization Power of TCR Antigenic Epitope Prediction Models

Model ReleasesDGX agent

arXiv:2606.04994v1 Announce Type: new Abstract: Accurate computational prediction of T cell receptor (TCR) antigen specificity would transform the study of T cell biology and enable scalable immune en

NVIDIA Nemotron 3 Ultra is on Fireworks, day zero. Nemotron Ultra is an open model for frontier reasoning and orchestration in long-running …

Model ReleasesDGX agent

NVIDIA Nemotron 3 Ultra is on Fireworks, day zero. Nemotron Ultra is an open model for frontier reasoning and orchestration in long-running autonomous agents. Think use cases like coding agents, deep

Proof-Carrying Agent Actions: Model-Agnostic Runtime Governance for Heterogeneous Agent Systems

Model ReleasesDGX agent

arXiv:2606.04104v1 Announce Type: cross Abstract: Agent systems execute through runtimes with very different control points: local coding tools, framework SDKs, managed agent platforms, API gateways,

Pulled the trigger today and switched 100% of Lindy traffic to DeepSeek v4, churning from Anthropic models. Saves us millions of $ and we're…

Model ReleasesDGX agent

Pulled the trigger today and switched 100% of Lindy traffic to DeepSeek v4, churning from Anthropic models. Saves us millions of $ and we're actually seeing an *increase* in performance on many core u

Robotics startup Generalist, which released its GEN-1 model to complete short physical tasks in April, raised 400M led by Radical Ventures at a 2B valuation (Dina Bass/Bloomberg)

Model ReleasesDGX agent

Dina Bass / Bloomberg: Robotics startup Generalist, which released its GEN-1 model to complete short physical tasks in April, raised 400M led by Radical Ventures at a 2B valuation — The company raised

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification

ResearchDGX agent

arXiv:2606.04579v1 Announce Type: new Abstract: While Process Reward Models (PRMs) have achieved remarkable success in mathematical reasoning, their application in complex scientific domains-such as b

That's a badass title, and it's true! Every day, it gets harder and harder to create tests that AI models can't beat. Reality is humanity's …

Model ReleasesDGX agent

That's a badass title, and it's true! Every day, it gets harder and harder to create tests that AI models can't beat. Reality is humanity's real last exam. Andon Labs' Real-World AI Evals: Claude call

Value Entanglement: Conflation Between Different Kinds of Good In (Some) Large Language Models

SafetyDGX agent

arXiv:2602.19101v2 Announce Type: replace-cross Abstract: Value alignment of Large Language Models (LLMs) requires us to empirically measure these models' actual, acquired representation of value. Amo

3 Jun 2026

Alibaba releases Qwen3.7-Plus, a multimodal proprietary model with a 1M-token context window, costing $2 per 1M tokens, 60% less than text-only Qwen3.7-Max (Carl Franzen/VentureBeat)

Model ReleasesDGX agent

Carl Franzen / VentureBeat: Alibaba releases Qwen3.7-Plus, a multimodal proprietary model with a 1M-token context window, costing $2 per 1M tokens, 60% less than text-only Qwen3.7-Max — However, like

Announcing Foundry Managed Compute: Run open models in Microsoft Foundry

HardwareDGX agent

Microsoft Foundry Managed Compute is a new GPU platform-as-a-service for hosting open-source and custom AI models behind the same endpoint, SDKs, and bill as frontier models. The post Announcing Found

AugMask: Training Diffusion Models on Incomplete Tabular Data via Stochastic Augmentation and Masking

ApplicationsDGX agent

arXiv:2606.03347v1 Announce Type: cross Abstract: Score-based diffusion models have emerged as prominent deep generative models; however, their application to tabular data remains challenging because

Bridging Auxiliary Constraints to Resolve Instruction Following in Large Reasoning Models

ResearchDGX agent

arXiv:2606.03624v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have demonstrated impressive capabilities in many tasks, yet they struggle with reliably following multiple instructions,

ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models

Model ReleasesDGX agent

arXiv:2606.03157v1 Announce Type: new Abstract: Large language models (LLMs) have been widely adopted in healthcare, yet they still encounter significant challenges in complex clinical decision-making

DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction

AgentsDGX agent

arXiv:2606.03874v1 Announce Type: new Abstract: We present DyaPlex, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of

Expert-Aware Causal Tracing of Factual Recall in Sparse MoE Language Models

Local AiDGX agent

arXiv:2606.03780v1 Announce Type: new Abstract: Causal tracing of factual recall has been studied predominantly in dense transformer language models, where interventions localize information flow to l

Frontier models are powerful advisors. On @harvey's Legal Agent Benchmark, a GLM 5.1 worker using Claude Opus 4.7 as a sparse advisor reache…

Model ReleasesDGX agent

Frontier models are powerful advisors. On @harvey's Legal Agent Benchmark, a GLM 5.1 worker using Claude Opus 4.7 as a sparse advisor reached 18/100 all-pass versus 14/100 for Opus alone, at 39% of th

Gemma 4 model load issues fixed in engine version 2.20.1. lms runtime update --all

Model ReleasesDGX agent

Gemma 4 model load issues fixed in engine version 2.20.1. lms runtime update --all Gemma 4 12B is here! Dense, mid-sized Gemma that fits right on your laptop - released by @google under Apache 2.0 Ava

Instant Personalized Large Language Model Adaptation via Hypernetwork

Model ReleasesDGX agent

arXiv:2510.16282v2 Announce Type: replace Abstract: Personalized large language models (LLMs) tailor content to individual preferences using user profiles or histories. However, existing parameter-eff

Large Language Models Are Overconfident in Their Own Responses

SafetyDGX agent

arXiv:2606.03437v1 Announce Type: new Abstract: Prior work has shown that instruction-tuned large language models (LLMs) are less well calibrated than their base pre-trained counterparts. However, lit

Measuring Weak-to-Strong Legibility of Reasoning Models

SafetyDGX agent

arXiv:2603.20508v2 Announce Type: replace-cross Abstract: Reasoning language models (RLMs) and the intermediate chains of thought they emit play an increasingly central role in multi-agent setups such

Model page: https://ollama.com/library/gemma4

Local AiDGX agent

Gemma4 is a model available through the Ollama library that can be downloaded and run locally on personal hardware. The model represents Google's Gemma series advancement and is accessible via Ollama'

Non-Identical Diffusion Models in MIMO-OFDM Channel Generation

ResearchDGX agent

arXiv:2509.01641v3 Announce Type: replace-cross Abstract: We propose a novel diffusion model, termed the non-identical diffusion model, and investigate its application to wireless orthogonal frequency

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation

SafetyDGX agent

arXiv:2606.03159v1 Announce Type: cross Abstract: As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. In closed-lo

Plan, Verify and Fill: A Structured Parallel Decoding Approach for Diffusion Language Models

Model ReleasesDGX agent

arXiv:2601.12247v3 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) present a promising non-sequential paradigm for text generation, distinct from standard autoregressive (AR) a

Position: Prioritize Identifying Structure, Not Complex Models, for Scientific Discovery

ResearchDGX agent

arXiv:2606.02632v1 Announce Type: cross Abstract: Modern Machine Learning (ML) and Artificial Intelligence (AI) models, especially large language models (LLMs), are increasingly used to generate scien

PyraMathBench: Evaluating and Improving Mathematical Capability in Large Language Models

Model ReleasesDGX agent

arXiv:2606.03858v1 Announce Type: new Abstract: Despite the pivotal role of numerical reasoning as the cornerstone of mathematical capabilities in large language models (LLMs) across applications, few

Startup discovery platform @harmonic_ai rebuilt Scout, their AI platform using Deep Agents and LangSmith. Deep Agents: One frontier model + …

Model ReleasesDGX agent

Startup discovery platform @harmonic_ai rebuilt Scout, their AI platform using Deep Agents and LangSmith. Deep Agents: One frontier model + two tool sets (global company data and firm-specific context

Text-to-Image Models Need Less from Text Encoders Than You Think

TutorialsDGX agent

arXiv:2606.03715v1 Announce Type: new Abstract: Text-to-image models rely on text prompts as their primary interface to human intent. Prompts are encoded by a text encoder into embeddings that conditi

Thinking Past the Answer: Evaluating Harmful Overthinking in Large Reasoning Models

ResearchDGX agent

arXiv:2606.02835v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by generating explicit intermediate reasoning traces through increased test-time compute, yet the assu

TimeOmni-VL: Unified Models for Time Series Understanding and Generation

ResearchDGX agent

arXiv:2602.17149v2 Announce Type: replace-cross Abstract: Recent time series modeling faces a sharp divide between numerical generation and semantic understanding, with research showing that generatio

Visual Graph Scaffolds for Structural Reasoning in Large Language Models

TutorialsDGX agent

arXiv:2606.02673v1 Announce Type: new Abstract: Graphs have been used to enhance large language models (LLMs) for structured reasoning, mostly as external knowledge sources are provided to models at t

We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-…

Model ReleasesDGX agent

We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-5.5’s agentic coding and tool use together with stronger int

When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models

SafetyDGX agent

arXiv:2606.03712v1 Announce Type: new Abstract: Graph Language Models (GLMs) have become a promising direction for adapting Large Language Models (LLMs) to graph learning tasks. By transforming graph

When Model Merging Breaks Routing: Training-Free Calibration for MoE

Model ReleasesDGX agent

arXiv:2606.03391v1 Announce Type: cross Abstract: Model merging has emerged as a cost-effective approach for consolidating the capabilities of multiple LLMs without retraining. However, existing mergi

← Previous
1…107108109110111…1009
Next →