AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,428 results
9 Jul 2026

[AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition

Model ReleasesDGX agent

SpaceX's AI division has launched Grok 4.5, positioned as their first Opus-class model following the acquisition of Cursor. The release represents a significant capability upgrade in SpaceX's AI offer

Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models

Model ReleasesDGX agent

arXiv:2607.07251v1 Announce Type: new Abstract: One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial rea

Inertia-1: An Open Exploration of Wearable Motion Foundation Models

ApplicationsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.06617v1 Announce Type: cross Abstract: Wearable motion sensing provides a continuous and scalable window into human behavior and health, making it a natural fit for foundation models, yet i

Just shared my favorite AI model and tool for every use case. Check out the list here: https://aiwithallie.beehiiv.com/p/the-best-ai-model-a…

Model ReleasesDGX agent

Just shared my favorite AI model and tool for every use case. Check out the list here: https://aiwithallie.beehiiv.com/p/the-best-ai-model-and-tool-for-every-use-case I'm writing a newsletter on my fa

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

Model ReleasesDGX agent

arXiv:2607.06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic

PrismML shrunk Qwen 3.6 to 27B parameters and ran it on an iPhone 17 Pro, bigger than any prior mobile model; sources: Apple talked with PrismML about its tech (Aaron Tilley/The Information)

Model ReleasesDGX agent

Aaron Tilley / The Information: PrismML shrunk Qwen 3.6 to 27B parameters and ran it on an iPhone 17 Pro, bigger than any prior mobile model; sources: Apple talked with PrismML about its tech — Apple

Residual-Conservative Model Predictive Path Integral Control

SafetyDGX agent

arXiv:2607.06950v1 Announce Type: cross Abstract: Sampling-based model predictive control methods handle nonlinear dynamics and complex cost landscapes through Monte Carlo rollouts, yet typically empl

SmartHomeSecure: Automated Detection and Repair of Smart Home Configuration Errors Using Large Language Models

Model ReleasesDGX agent

arXiv:2607.06748v1 Announce Type: cross Abstract: Smart home automation platforms increasingly rely on user-authored YAML configuration files to define device behaviors, but these files are prone to s

spent a ton of time working on this!! so happy to see the results come out. very impressive model from @NVIDIAAI that we fit our harness to …

Model ReleasesDGX agent

spent a ton of time working on this!! so happy to see the results come out. very impressive model from @NVIDIAAI that we fit our harness to so we could get better performance! We tuned the harness for

Stability of Flow Models for Graph Signals

ApplicationsDGX agent

arXiv:2607.07510v1 Announce Type: cross Abstract: Generating signals on graphs requires permutation-equivariant models that exhibit stability with respect to relative structural perturbations. While f

The Rank-One Corner: How Much Value Equivalence Does a Task Need from a World Model?

ResearchDGX agent

arXiv:2607.06640v1 Announce Type: cross Abstract: A learned world model is usually judged by how faithfully it reconstructs its observations or predicts reward, as though quality were something the mo

8 Jul 2026

A Definition and Roadmap for World Models

TutorialsDGX agent

arXiv:2607.06401v1 Announce Type: new Abstract: World models -- internal simulators that learn the structure and dynamics of an environment -- have become one of the most actively debated concepts in

As coding models improve, evals need to become harder, fairer, and more trustworthy. Better benchmarks help the field understand real progre…

Model ReleasesDGX agent

As AI coding models advance in capability, current evaluation benchmarks must evolve to remain challenging and meaningful measures of progress. OpenAI argues that improved benchmarks need to be harder

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models

Model ReleasesDGX agent

arXiv:2607.06534v1 Announce Type: new Abstract: Existing 3D scene-grounded Large Language Models (3D-LLMs) focus on answering questions grounded in simplified single-room 3D scenes, lacking the abilit

Claude power users: 'Fable 5 is the best' Codex power users: 'GPT-5.6 is the best' Reality: Loyalty to a single model provider is a terrible…

Model ReleasesDGX agent

Claude power users: 'Fable 5 is the best' Codex power users: 'GPT-5.6 is the best' Reality: Loyalty to a single model provider is a terrible strategy. The smart choice: clever orchestration between fr

Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

Model ReleasesDGX agent

arXiv:2607.06482v1 Announce Type: cross Abstract: Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fac

Mistral launches Robostral Navigate, a hardware-agnostic robotics navigation model trained via simulation that uses a single camera and basic language prompts (Benoit Berthelot/Bloomberg)

Model ReleasesDGX agent

Benoit Berthelot / Bloomberg: Mistral launches Robostral Navigate, a hardware-agnostic robotics navigation model trained via simulation that uses a single camera and basic language prompts — Mistral A

Multi-Teacher Contrastive Distillation for Edge-Efficient Pathology Foundation Models

Model ReleasesDGX agent

arXiv:2607.05533v1 Announce Type: new Abstract: Computational pathology foundation models (PFMs) have advanced whole-slide image analysis. However, their size and inference cost hinder local deploymen

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

Model ReleasesDGX agent

arXiv:2607.05722v1 Announce Type: new Abstract: We introduce Nemotron-Labs-Diffusion, a tri-mode language model (LM) that unifies AR, diffusion, and self-speculation decoding within a single architect

SpaceXAI launches Grok 4.5, its first model built in partnership with Cursor, designed to 'handle difficult, long-running' legal, finance, and coding tasks (Carmen Arroyo/Bloomberg)

Model ReleasesDGX agent

Carmen Arroyo / Bloomberg: SpaceXAI launches Grok 4.5, its first model built in partnership with Cursor, designed to “handle difficult, long-running” legal, finance, and coding tasks — SpaceXAI has un

SpaceXAI’s newest AI model Grok 4.5 dramatically undercuts Anthropic and OpenAI on price

Model ReleasesDGX agent

Elon Musk’s SpaceXAI Corp. has released a new model called Grok 4.5, in what is its first major launch since it went public a few weeks earlier. In a blog post earlier today, the company said Grok 4.5

SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation

Model ReleasesDGX agent

arXiv:2607.05721v1 Announce Type: new Abstract: Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Model ReleasesDGX agent

arXiv:2502.09696v3 Announce Type: replace Abstract: Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children

7 Jul 2026

A Fair Benchmarking of Deep Relational Database Learning Models

Model ReleasesDGX agent

arXiv:2607.03659v1 Announce Type: cross Abstract: Relational databases (RDBs) are the primary data infrastructure in many enterprises, yet recent deep learning methods designed for RDBs have been eval

A Technical Survey of Reinforcement Learning Techniques for Large Language Models

Model ReleasesDGX agent

arXiv:2507.04136v2 Announce Type: replace Abstract: This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Poli

A Unified Framework for In-Context Learning with Causal and Masked Language Models

ResearchDGX agent

arXiv:2607.04081v1 Announce Type: new Abstract: In-context learning (ICL) has emerged as a central capability of pretrained language models, yet its theoretical analysis has focused primarily on causa

Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human-AI Collaboration

Model ReleasesDGX agent

arXiv:2607.02563v1 Announce Type: cross Abstract: Diffusion-based text-to-image models can synthesize complex and highly structured visual content, yet the emergence and evolution of semantic structur

Banger compression paper from NVIDIA. (bookmark it) Bigger MoE models keep winning on quality, but serving them at interactive latency is st…

Model ReleasesDGX agent

Banger compression paper from NVIDIA. (bookmark it) Bigger MoE models keep winning on quality, but serving them at interactive latency is still hard. NVIDIA compresses the hybrid MoE Nemotron-3-Super

CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection

Model ReleasesDGX agent

arXiv:2607.02930v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel in diverse vision tasks, but full-parameter retraining is computationally expensive as real-world knowled

Do Medical Vision Language Models Actually See? A Counterfactual Grounding Framework and Hard-Negative Contrastive Training for Visually-Reliant Medical VLMs

Model ReleasesDGX agent

arXiv:2607.03647v1 Announce Type: new Abstract: Large vision language models (VLMs) report strong accuracy on medical question-answering, yet it remains unclear whether they reason from visual evidenc

DREAMSTEER: Latent World Models Can Steer VLA Policies During Deployment Without Any Finetuning

Model ReleasesDGX agent

arXiv:2607.02865v1 Announce Type: new Abstract: Pretrained vision-language-action (VLA) policies show promising zero-shot generalization, but often fail under deployment-time distribution shift, leadi

DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics

ResearchDGX agent

arXiv:2607.04112v1 Announce Type: cross Abstract: Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs require model

Enhancing the Forecasting Capability of Multi-Model Blending Algorithms for Extreme Precipitation via Joint Use of Station and Gridded Observations

SafetyDGX agent

arXiv:2607.04862v1 Announce Type: new Abstract: Accurate extreme precipitation forecasting is critical for disaster mitigation but remains challenging for numerical weather prediction (NWP) models due

Erasing Without Collateral Damage: Precise Concept Removal in Diffusion Models

Model ReleasesDGX agent

arXiv:2607.05274v1 Announce Type: new Abstract: Training-free concept erasure is an attractive mechanism for controlling text-to-image diffusion models, but precise erasure often comes at the cost of

Excellent investigation of the mechanics of language models. 🫢A notable blindspot, however: several teams have actually been **directly** c…

Model ReleasesDGX agent

Excellent investigation of the mechanics of language models. 🫢A notable blindspot, however: several teams have actually been **directly** comparing the working of LLMs to those of the human brain 🧠 fo

Exploring the Rashomon Set for Concept-Based Models

ResearchDGX agent

arXiv:2511.19636v2 Announce Type: replace-cross Abstract: In many machine learning problems, there may exist multiple models that achieve nearly identical predictive performance while relying on funda

Framework for Grouping Local Process Models

Local AiDGX agent

arXiv:2607.04856v1 Announce Type: new Abstract: Local Process Models (LPMs) are an underexplored concept in process mining. LPMs describe patterns in event data considering sequence, choice, concurren

Geographic Diversity Beats Data Volume for Cross-Domain Generalization in Zero-Label JEPA Driving World Models

ResearchDGX agent

arXiv:2607.04500v1 Announce Type: new Abstract: Self-supervised latent world models can assign a surprise score to driving scenarios without any human labels. A natural follow-up question is whether s

Geometry of Ordinal Representations in Language Models

Model ReleasesDGX agent

arXiv:2607.04167v1 Announce Type: new Abstract: Recent work showed that language models represent character counts on curved 1D manifolds, with attention heads performing geometric transformations to

How Utilitarian Are OpenAI's Models Really? Replicating and Reinterpreting Pfeffer, Krugel, and Uhl (2025)

SafetyDGX agent

arXiv:2603.22730v2 Announce Type: replace Abstract: Pfeffer, Krugel, and Uhl (2025) report that OpenAI's reasoning model o1-mini produces more utilitarian responses to the trolley problem and footbrid

Integrating Neural Encoders in Bayesian Generalized Linear Mixed Models for Multimodal Data

Model ReleasesDGX agent

arXiv:2607.04647v1 Announce Type: cross Abstract: Scalable Bayesian inference for generalized linear mixed models (GLMMs) provides uncertainty-aware analysis of correlated longitudinal data, but exist

Modular Foundation Models for Time-Series Perception in Digital Twins

Model ReleasesDGX agent

arXiv:2607.03585v1 Announce Type: new Abstract: Engineering Digital Twins and Prognostics and Health Management (PHM) systems rely on robust perception modules to extract actionable information from h

OmniFocus: Query-Guided Modality-Balanced Token Compression for Omni-Modal Large Language Models

Model ReleasesDGX agent

arXiv:2607.03050v1 Announce Type: cross Abstract: Omni modal large language models (OmniLLMs) have attracted wide attention for their ability to jointly process audio and video, but they generate larg

Revealing Hidden Model Behaviors with Task-Specific Self-Reports

TutorialsDGX agent

arXiv:2607.03640v1 Announce Type: cross Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt

Signal or Noise? Understanding Generative Models for Real-World Sensor Time Series

ApplicationsDGX agent

arXiv:2607.04245v1 Announce Type: cross Abstract: Generative models have changed how machine learning represents complex data distributions, especially in language and vision, yet many real-world syst

Stacked LoRA for Subject-Adaptive EEG Foundation Models in Motor Imagery Decoding

Model ReleasesDGX agent

arXiv:2607.03094v1 Announce Type: new Abstract: Electroencephalography (EEG) decoding for brain-computer interfaces (BCIs) faces a major challenge: substantial inter-subject variability limits effecti

Training Hybrid Block Diffusion Language Models with Partial Bidirectionality

Model ReleasesDGX agent

arXiv:2607.02805v1 Announce Type: cross Abstract: High-throughput long-context generation is one of the central challenges for large language models. Generation is typically memory-bandwidth-bound rat

UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation

AgentsDGX agent

arXiv:2607.05133v1 Announce Type: new Abstract: World Action Models (WAMs) have shown strong potential for improving action generalization in autonomous driving by using future video prediction as den

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

Model ReleasesDGX agent

arXiv:2510.11098v5 Announce Type: replace-cross Abstract: Recent advances in large audio language models (LALMs) have greatly enhanced multimodal conversational systems. However, existing benchmarks r

VISTA: Auditing Semantic Divergence in Vision-Language Models

ResearchDGX agent

arXiv:2607.02995v1 Announce Type: cross Abstract: Vision-language models can exhibit visual concept-conditioned divergence: given images containing demographic features, corporate logos, or ideologica

WAM4D: Fast 4D World Action Model via Spatial Register Tokens

ApplicationsDGX agent

arXiv:2606.14048v2 Announce Type: replace Abstract: World action models (WAMs) have recently shown promise in jointly modeling future observations and executable robot actions. However, most existing

Which Algorithm Specification Formats Help Language Models Implement Machine Learning Algorithms?

Model ReleasesDGX agent

arXiv:2607.03158v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to implement algorithms from research manuscripts, but papers often leave implementation choices im

Wrong Before Right: Late Rescue and Interface Failure in Aligned Language Models

Model ReleasesDGX agent

arXiv:2607.04640v1 Announce Type: new Abstract: We study how correctness is assembled inside aligned language models, not only whether the final answer is right. Using layer-wise difference-in-differe

6 Jul 2026

Claude Opus 4.8 and Sonnet 5 seem worse at tool calls than older models, likely due to post-training that assumes Claude Code-like harnesses as targets (Armin Ronacher/Armin Ronacher's Thoughts and Writings)

Model ReleasesDGX agent

Armin Ronacher / Armin Ronacher's Thoughts and Writings: Claude Opus 4.8 and Sonnet 5 seem worse at tool calls than older models, likely due to post-training that assumes Claude Code-like harnesses as

3 Jul 2026

GLM-5.2 is now selectable in Claude Code via Hugging Face🤗 Inference Providers + hf-claude. Open models are becoming easier to plug directl…

Model ReleasesDGX agent

GLM-5.2, an open-source model available through Hugging Face, can now be selected and used within Claude Code through Hugging Face Inference Providers and the hf-claude integration. This development d

Interpreting Global Perturbation Robustness of Image Models using Axiomatic Spectral Importance Decomposition

ResearchDGX agent

arXiv:2408.01139v4 Announce Type: replace Abstract: Perturbation robustness evaluates the vulnerabilities of models, arising from a variety of perturbations, such as data corruptions and adversarial a

Meta to release new AI model with advanced coding capabilities ‘soon’

Model ReleasesDGX agent

Meta Platforms Inc. is gearing up to release a new version of its flagship Muse Spark artificial intelligence model. Alexandr Wang, the company’s chief AI officer, wrote on X today that the update wil

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

Model ReleasesDGX agent

arXiv:2607.01813v1 Announce Type: cross Abstract: Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to

OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models

ResearchDGX agent

arXiv:2607.01977v1 Announce Type: new Abstract: Ontology learning (OL) aims to automatically construct structured knowledge models from text, yet progress remains fragmented across methods, domains, a

Prompt engineering is costing you money. Learn how to fine-tune your models. Stuffing a lot of text into every single API call slows down yo…

AgentsDGX agent

Prompt engineering is costing you money. Learn how to fine-tune your models. Stuffing a lot of text into every single API call slows down your app because the model has to process all those tokens bef

← Previous
1…8081828384…1008
Next →