AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,929 results
Model Releases

In-Context World Modeling for Robotic Control

DGX agent

arXiv:2606.26025v1 Announce Type: cross Abstract: Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoints or robot morphologies, because

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

DGX agent

arXiv:2606.25984v1 Announce Type: cross Abstract: Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately reconstruct and

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models

DGX agent

arXiv:2604.05738v2 Announce Type: replace Abstract: Medical Vision-Language Models (Med-VLMs) have achieved expert-level proficiency in interpreting diagnostic imaging. However, current models are pre

model-releasesarxiv-cs-cl
25 Jun 2026
Model Releases

Model-agnostic Mitigation Strategies of Data Imbalance for Regression

DGX agent

arXiv:2506.01486v2 Announce Type: replace Abstract: Data imbalance persists as a pervasive challenge in regression tasks, introducing bias in model performance and undermining predictive reliability.

model-releasesarxiv-cs-lg
25 Jun 2026
Model Releases

Steering Vision-Language Models with Joint Sparse Autoencoders

DGX agent

arXiv:2606.25657v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) have shown promise for analyzing language models, but applying them to vision-language models (VLMs) often yields representat

model-releasesarxiv-cs-cv
25 Jun 2026
Model Releases

ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation

DGX agent

arXiv:2606.23835v1 Announce Type: new Abstract: ABACUS is a unified vision-language model that handles object counting, crowd counting, referring-expression counting, and count-faithful image generati

model-releasesarxiv-cs-cv
24 Jun 2026
Model Releases

Gemini 3 Pro was the first model to achieve at least 23% on ARC-AGI-2, which it did in November, 2025 (it actually scored 31%). So the 8-12 …

DGX agent

Gemini 3 Pro was the first model to achieve at least 23% on ARC-AGI-2, which it did in November, 2025 (it actually scored 31%). So the 8-12 month gap between closed and open weights models still seems

model-releasesethan-mollick--x
24 Jun 2026
Model Releases

HANCLIP: A Family of Hyperbolic Angular Negation Vision Language Models

DGX agent

arXiv:2606.23843v1 Announce Type: new Abstract: Vision-Language Models (VLMs) are typically pre-trained on large-scale image-text datasets to capture semantic correspondences between visual content an

model-releasesarxiv-cs-cv
24 Jun 2026
Safety

Quant Convergence: Bridging Classical Value Investing and Modern Factor Models for Systematic Equity Selection

DGX agent

arXiv:2606.24575v1 Announce Type: new Abstract: Modern finance relies heavily on complex machine learning models to find patterns in the stock market. However, as these AI models get more complicated,

safetyarxiv-cs-ai
24 Jun 2026
Model Releases

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning

DGX agent

arXiv:2412.08147v2 Announce Type: replace-cross Abstract: Pareto fronts are useful to find good task-mixing strategies for multitask finetuning, but they are also costly to compute. To reduce costs, r

model-releasesarxiv-cs-ai
24 Jun 2026
Research

2D Versus 3D Diffusion for In Silico Training of Interventional X-ray AI Models

DGX agent

arXiv:2606.21414v1 Announce Type: cross Abstract: The ability to synthesize realistic X-ray images has catalyzed the development of AI models for X-ray image-guided procedures, which otherwise suffer

researcharxiv-cs-cv
23 Jun 2026
Model Releases

Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think

DGX agent

arXiv:2606.20246v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models pre-trained on massive video-robot datasets have revolutionized robotic manipulation, yet their multi-billion pa

model-releasesarxiv-cs-ro
23 Jun 2026
Model Releases

Generating Fearful Images: Investigating Potential Emotional Biases in Image-Generation Models

DGX agent

arXiv:2411.05985v3 Announce Type: replace-cross Abstract: This paper examines potential biases and inconsistencies in the emotions evoked by images produced by generative artificial intelligence (AI)

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Holo-World: Unified Camera, Object and Weather Control for Video World Model

DGX agent

arXiv:2606.20083v2 Announce Type: replace Abstract: Video world models are moving toward preserving an observed world under controllable camera and object motion while allowing its environmental state

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

MapReason-OSM: Can Vision-Language Models Make Graph-Verifiable Mobility Decisions from Street Maps ?

DGX agent

arXiv:2606.22597v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used to read maps for logistics, delivery, and accessible navigation, where the output is an actionable d

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Right Knowledge, Wrong Answer: Test-Time Steering for Temporal Fact Conflicts in Open-Weight Language Models

DGX agent

arXiv:2606.20959v1 Announce Type: new Abstract: Large language models can store both outdated facts and newer superseding facts in their parameters, but standard prompting may still elicit the outdate

model-releasesarxiv-cs-lg
23 Jun 2026
Model Releases

SegTME-UNI2: A Foundation Model-Based Framework for Generalisable Multiclass Cell Segmentation and LLM-Driven Tumour Microenvironment Characterisation in Histopathology

DGX agent

arXiv:2606.17702v2 Announce Type: replace Abstract: Characterising the tumour microenvironment (TME) from routine H&E-stained histology images requires simultaneous cell segmentation, feature extracti

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

SparseWorld: Enhancing End-to-End Autonomous Driving via World Models with Sparse Scene Representation

DGX agent

arXiv:2605.24354v2 Announce Type: replace Abstract: Recently, world models have made significant progress in enhancing end-to-end driving systems through both future situation forecasting and improved

model-releasesarxiv-cs-cv
23 Jun 2026
Model Releases

Tapered Language Models

DGX agent

arXiv:2606.23670v1 Announce Type: new Abstract: Modern language models, including transformer, recurrent, and memory-based variants, share a common chassis: a stack of identical layers in which parame

model-releasesarxiv-cs-lg
23 Jun 2026
Research

World Action Models: A Survey

DGX agent

arXiv:2606.20781v1 Announce Type: cross Abstract: World Action Models (WAMs) are embodied predictive-action models that make a forecast of the future available to action. Recent WAMs repurpose large v

researcharxiv-cs-cv
23 Jun 2026
Model Releases

GLM-5.2 leads open weights models and sits at #3 overall on GDPval-AA, a real-world agentic work benchmark GLM-5.2 from @Zai_org scores 1524…

DGX agent

GLM-5.2 leads open weights models and sits at #3 overall on GDPval-AA, a real-world agentic work benchmark GLM-5.2 from @Zai_org scores 1524 Elo on GDPval-AA, which measures performance on real-world,

model-releasesclem-delangue--x
22 Jun 2026
Model Releases

That's right, an open, MIT-licensed model beating GPT-5.5 (xhigh) on real-world agentic work! 🔥 Available for free on @huggingface for anyo…

DGX agent

That's right, an open, MIT-licensed model beating GPT-5.5 (xhigh) on real-world agentic work! 🔥 Available for free on @huggingface for anyone to build on top off GLM-5.2 leads open weights models and

model-releasesclem-delangue--x
22 Jun 2026
Research

Fast Speech Foundation Model Distillation Using Interleaved Stacking

DGX agent

arXiv:2606.11766v1 Announce Type: cross Abstract: Distilling a large speech foundation model (SFM) into an efficient student model has been successfully applied to low-resource environments. Although

researcharxiv-cs-ai
11 Jun 2026
Model Releases

Improving Cross-Format Robustness in Language Models with Multi-Format Training

DGX agent

arXiv:2606.11643v1 Announce Type: new Abstract: Large language models often remain sensitive to answer format: a question solved correctly in one form may fail in another semantically equivalent form.

model-releasesarxiv-cs-cl
11 Jun 2026
Safety

Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models

DGX agent

arXiv:2606.11409v1 Announce Type: cross Abstract: Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly tr

safetyarxiv-cs-ai
11 Jun 2026
Model Releases

The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

DGX agent

arXiv:2606.12289v1 Announce Type: cross Abstract: As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and controlling

model-releasesarxiv-cs-ai
11 Jun 2026
Research

Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View

DGX agent

arXiv:2603.05573v2 Announce Type: replace Abstract: Scalable sequence models, such as Transformer variants and structured state-space models, often trade expressivity power for sequence-level parallel

researcharxiv-cs-lg
11 Jun 2026
Model Releases

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

DGX agent

arXiv:2606.11063v1 Announce Type: new Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This part

model-releasesarxiv-cs-ai
10 Jun 2026
Model Releases

Divide-and-Conquer Modeling for the CTF-4-Science Lorenz Benchmark

DGX agent

arXiv:2606.10084v1 Announce Type: cross Abstract: This work presents a divide-and-conquer modeling strategy for the CTF-4-Science Lorenz benchmark, which evaluates chaotic-system prediction across twe

model-releasesarxiv-cs-ai
10 Jun 2026
Safety

MARCH: Model-Assisted Reinforcement Learning for the Perceptive Control of Humanoids over Sparse Footholds

DGX agent

arXiv:2606.10288v1 Announce Type: new Abstract: Perceptive bipedal locomotion over sparse terrain remains a difficult challenge: model-based methods are precise but brittle to uncertainty, while model

safetyarxiv-cs-ro
10 Jun 2026
Model Releases

Parallel Causal Associative Fields: Gated Sparse Memory for Long-Context Language Modeling

DGX agent

arXiv:2606.10435v1 Announce Type: cross Abstract: Transformers achieve strong language modeling performance by providing direct token-to-token communication paths, but causal self-attention scales qua

model-releasesarxiv-cs-cl
10 Jun 2026
Safety

When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models

DGX agent

arXiv:2512.06343v3 Announce Type: replace-cross Abstract: Reward models are central to Large Language Model (LLM) alignment within the framework of RLHF. The standard objective used in reward modeling

safetyarxiv-cs-ai
10 Jun 2026
Model Releases

WorldOlympiad: Can Your World Model Survive a Triathlon?

DGX agent

arXiv:2606.11129v1 Announce Type: new Abstract: We introduce WorldOlympiad, a benchmark for diagnosing video-based world models across physical faithfulness, geometric consistency, and interaction fid

model-releasesarxiv-cs-cv
10 Jun 2026
Safety

AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing

DGX agent

arXiv:2606.09811v1 Announce Type: cross Abstract: World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

BREAKING: Anthropic just dropped Claude Fable 5—this is Mythos, made safe for public release. It is the best coding model in the world. We'v…

DGX agent

BREAKING: Anthropic just dropped Claude Fable 5—this is Mythos, made safe for public release. It is the best coding model in the world. We've been testing it internally @every for the last week or so

model-releasessimon-willison--x
9 Jun 2026
Model Releases

CoVEBench: Can Video Editing Models Handle Complex Instructions?

DGX agent

arXiv:2606.08415v1 Announce Type: cross Abstract: While recent text-guided video editing models excel at elementary tasks (e.g., style transfer, object insertion), real-world user requests are highly

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Domain-Adapted Small Language Models with Hybrid Post-Processing: Achieving Cost-Efficient, Low-Latency Multi-Label Structured Prediction via LoRA Fine-Tuning on Scarce Data

DGX agent

arXiv:2606.05781v2 Announce Type: replace Abstract: Deploying frontier large language models (LLMs) for domain-specific structured evaluation tasks incurs prohibitive latency, cost, and data-privacy o

model-releasesarxiv-cs-lg
9 Jun 2026
Model Releases

Executable World Models for ARC-AGI-3 in the Era of Coding Agents

DGX agent

arXiv:2605.05138v2 Announce Type: replace Abstract: We evaluate an initial coding-agent system for ARC-AGI-3 in which the agent maintains an executable Python world model, verifies it against previous

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

In-Context Learning of Temporal Point Processes with Foundation Inference Models

DGX agent

arXiv:2509.24762v3 Announce Type: replace Abstract: Modeling event sequences of multiple event types with marked temporal point processes (MTPPs) provides a principled way to uncover governing dynamic

model-releasesarxiv-cs-lg
9 Jun 2026
Tools

Introducing North Mini Code: Cohere’s First Model For Developers

DGX agent

Cohere Labs introduced North Mini Code, a compact code generation model designed specifically for developers, marking Cohere's entry into specialized coding models. The model is optimized for efficien

toolshugging-face
9 Jun 2026
Safety

MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models

DGX agent

arXiv:2606.07996v1 Announce Type: cross Abstract: Pretraining is fundamental to the development of Large Language Models (LLMs), yet the opacity of pretraining data complicates model analysis and rais

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Prescriptive Scaling Reveals the Evolution of Language Model Capabilities

DGX agent

arXiv:2602.15327v2 Announce Type: replace-cross Abstract: Machine learning model performance improvements tend to arise from competition and application. For deployment, we consider prescriptive scali

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

Readable Yet Unpredictable: Rotated-Outcome Prediction in Vision-Language Models

DGX agent

arXiv:2606.07641v1 Announce Type: new Abstract: Can vision-language models predict what a 180{eg} rotation would reveal from the original image alone? We study this ability through Rotated-Outcome Pre

model-releasesarxiv-cs-cv
9 Jun 2026
Model Releases

Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models

DGX agent

arXiv:2603.19183v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have emerged as a promising approach for general-purpose robot manipulation. However, little research has mechan

model-releasesarxiv-cs-ro
9 Jun 2026
Model Releases

This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great…

DGX agent

This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *

model-releaseskarpathy--x
9 Jun 2026
Applications

UAOR: Uncertainty-aware Observation Reinjection for Vision-Language-Action Models

DGX agent

arXiv:2602.18020v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models leverage pretrained Vision-Language Models (VLMs) as backbones to map images and instructions to actions, demons

applicationsarxiv-cs-cv
9 Jun 2026
Safety

Unifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic Tasks

DGX agent

arXiv:2606.08775v1 Announce Type: cross Abstract: Visual world models have shown great potential in learning complex system dynamics. Recent advancements leverage these models as transition functions

safetyarxiv-cs-ai
9 Jun 2026
Model Releases

Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory

DGX agent

arXiv:2606.06523v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence. Despite recen

model-releasesarxiv-cs-ai
8 Jun 2026
← Previous
1…6162636465…1249
Next →