AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,490 results
9 Jun 2026

Multimodal Generative Engine Optimization: Rank Manipulation for Vision-Language Model Rankers

SafetyDGX agent

arXiv:2601.12263v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) integrate visual and textual knowledge into unified representations that increasingly underpin modern retrieval

Multimodal Large Language Models as Synthetic Participants in Video-Based Studies: An Evaluation

Model ReleasesDGX agent

arXiv:2606.07541v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have shown strong performance on objective tasks such as video understanding and reasoning. However, it remai

Phantom transitions in language model fine-tuning

Model ReleasesDGX agent

arXiv:2606.07559v1 Announce Type: cross Abstract: Fine-tuning a language model on contexts whose correct completion has a near-synonym competitor often fails silently. The cross-entropy loss decreases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction

Model ReleasesDGX agent

arXiv:2606.09788v1 Announce Type: new Abstract: Large-scale document processing requires contextually aware table extraction (TE) that is both accurate and efficient. Yet current approaches require bi

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation

ResearchDGX agent

arXiv:2606.08679v1 Announce Type: cross Abstract: Pretrained models are often evaluated on multi-task leaderboards to measure their applicability in diverse contexts. However, current methods for aggr

Scaling by Diversified Experience for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.09009v1 Announce Type: new Abstract: Vision-Language-Action models face significant challenges in real-world deployment due to the entanglement of high-level reasoning with low-level contro

See how Claude Fable 5 compares across every model: http://cursor.com/evals

Model ReleasesDGX agent

Claude Fable 5 is compared against other AI models on various evaluation metrics through Cursor's benchmarking tool. The evaluation likely covers performance across different tasks such as coding, rea

Should Demand Models Incorporate Competitor Prices? Oblivious Learning and Algorithmic Collusion

ResearchDGX agent

arXiv:2606.05363v2 Announce Type: replace-cross Abstract: On a platform with many sellers, should a pricing algorithm explicitly model competitors' prices when learning demand? Classical learning argu

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.07861v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored.

Unified Energy for Invariant and Independent Decoding in Diffusion Language Models

ResearchDGX agent

arXiv:2606.09159v1 Announce Type: cross Abstract: Diffusion Language Models (DLMs) enable parallel text generation by iteratively denoising a full sequence, offering attractive flexibility compared to

We talk a lot about how important it is to set up self-verification loops. Especially in the age of powerful models that can run for long pe…

Model ReleasesDGX agent

We talk a lot about how important it is to set up self-verification loops. Especially in the age of powerful models that can run for long periods of time, self-verification is a key ingredient that en

8 Jun 2026

ActiveGrasp: Information-Guided Active Grasping with Calibrated Energy-based Model

Model ReleasesDGX agent

arXiv:2511.12795v2 Announce Type: replace Abstract: Grasping in a densely cluttered environment is a challenging task for robots. Previous methods tried to solve this problem by actively gathering mul

Beyond Rubrics: Exploration-Guided Evaluation Skills for Reward Modeling

Model ReleasesDGX agent

arXiv:2606.07040v1 Announce Type: new Abstract: Open-ended reward modeling requires judges that can follow subtle, domain-specific preferences when verifiable answers are unavailable. Existing rubric-

Diagnosing Visual Ignorance in Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.06890v1 Announce Type: new Abstract: Vision-Language Models (VLMs) frequently rely on language priors, producing confident answers that are weakly grounded in visual evidence. While this be

Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios

SafetyDGX agent

arXiv:2606.06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know. Existing benchmarks emphasize domain

Good take My guess is - demand for intelligence is near infinite - but 80% of workloads will be running on 99% cheaper models within 12-18 m…

HardwareDGX agent

Good take My guess is - demand for intelligence is near infinite - but 80% of workloads will be running on 99% cheaper models within 12-18 months - 20% of workloads will still run on latest gen models

GuideCAD: A Lightweight Multimodal Framework for 3D CAD Model Generation via Prefix Embedding

Model ReleasesDGX agent

arXiv:2606.07024v1 Announce Type: new Abstract: Multi-modal approaches used for 3D CAD generation require substantial computational resources, necessitating efficient training. To address this, we pro

7 Jun 2026

Have been extensively testing Claude Workflows this weekend, with the best model possible. Threw it at my whole code base, combing for bugs.…

Model ReleasesDGX agent

Have been extensively testing Claude Workflows this weekend, with the best model possible. Threw it at my whole code base, combing for bugs. 144 found and fixed! Geez... It is a large code base, for s

6 Jun 2026

Finite Element-Based Material Learning via Automatic Differentiation: Learning constitutive neural network models from full-field deformation data

ResearchDGX agent

arXiv:2606.05199v1 Announce Type: cross Abstract: The identification of constitutive neural network models from heterogeneous full-field deformation data provides a robust alternative to traditional c

LatentWave: JEPA Pretraining for Wireless Foundation Models

SafetyDGX agent

arXiv:2606.06373v1 Announce Type: cross Abstract: Wireless foundation models have emerged as a promising alternative to building separate models for each wireless task. However, existing approaches re

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models

Model ReleasesDGX agent

arXiv:2606.05429v1 Announce Type: new Abstract: Post-training quantization (PTQ) is critical for the efficient deployment of large language models (LLMs). Recent ultra-low-bit PTQ methods rely on rigi

The Gemini Pro models do not seem to be iterating anywhere near as quickly as Claude or GPT (last release was 3.1 Pro in February). Its caus…

Model ReleasesDGX agent

The Gemini Pro models do not seem to be iterating anywhere near as quickly as Claude or GPT (last release was 3.1 Pro in February). Its causing a growing performance gap between Google and the other t

Towards Unified and Data-Efficient Prognostics and Health Management with Tabular Foundation Models

ResearchDGX agent

arXiv:2606.05481v1 Announce Type: cross Abstract: Data-driven Prognostics and Health Management (PHM) uses time-varying condition-monitoring data to diagnose system states and estimate remaining usefu

5 Jun 2026

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding

Local AiDGX agent

arXiv:2606.06155v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models leverage the rich world knowledge of pretrained vision-language models (VLMs) to enable instruction-following robo

Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.05531v1 Announce Type: cross Abstract: Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and ch

Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolution

Model ReleasesDGX agent

arXiv:2606.06492v1 Announce Type: cross Abstract: Code language models need repository-level context to resolve imports, APIs, and project conventions. Existing methods inject this knowledge as long i

Gemma 4 Quantization-Aware Training (QAT) weights are now available on Ollama! They reduce memory requirements while maintaining model quali…

Model ReleasesDGX agent

Gemma 4 Quantization-Aware Training (QAT) weights are now available on Ollama! They reduce memory requirements while maintaining model quality. E2B: ollama run gemma4:e2b-it-qat E4B: ollama run gemma4

MASF: A Multi-Model Adaptive Selection Framework for Abstractive Text summarization

ResearchDGX agent

arXiv:2606.05494v1 Announce Type: new Abstract: Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents a Multi-Model

Merging model-based control with multi-agent reinforcement learning for multi-agent cooperative teaming strategies

AgentsDGX agent

arXiv:2606.06011v1 Announce Type: new Abstract: In this work, we propose a framework that combines multi-agent reinforcement learning (MARL) with model-based control to achieve safe, dynamically feasi

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.05744v1 Announce Type: new Abstract: Spatial planning maps are central to territorial governance, translating planning objectives, regulations, and spatial strategies into visual forms for

Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion

Model ReleasesDGX agent

arXiv:2510.22768v2 Announce Type: replace Abstract: As autonomous agents increasingly interact, they inevitably attempt to influence one another. While prior work in text-only settings has explored th

Thousand Token Wood: shipping a multi-agent economy on a 3B model

AgentsDGX agent

Thousand Token Wood is a multi-agent economy simulation built on a 3 billion parameter language model, demonstrating how small models can power complex interactive systems with multiple agents. The pr

What are the most capable LLM models I can run on my laptop?

Local AiDGX agent

A discussion on r/ollama exploring which high-performance LLM models can be effectively run locally on standard laptop hardware , likely covering model size comparisons, hardware requirements, and per

4 Jun 2026

AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?

Model ReleasesDGX agent

arXiv:2606.05080v1 Announce Type: new Abstract: Scientific and engineering progress is fundamentally a long-horizon iterative process: proposing changes, running experiments, measuring outcomes, and c

Beyond Objective Equivalence: Constraint Injection for LLM-Based Optimization Modeling on Vehicle Routing Problems

Model ReleasesDGX agent

arXiv:2606.04816v1 Announce Type: new Abstract: Large language models (LLMs) increasingly translate natural-language optimization problems into executable solver code. Yet for constraint-dense operati

Data Attribution in Large Language Models via Bidirectional Gradient Optimization

ResearchDGX agent

arXiv:2606.04928v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed across diverse applications, raising critical questions for governance, accountability, and dat

Effective vocabulary expansion of multilingual language models for extremely low-resource languages

TutorialsDGX agent

arXiv:2602.09388v2 Announce Type: replace Abstract: Multilingual pre-trained language models(mPLMs) offer significant benefits for many low-resource languages. To further expand the range of languages

Gradient estimators for parameter inference in discrete stochastic kinetic models

Model ReleasesDGX agent

arXiv:2604.02121v2 Announce Type: replace-cross Abstract: Stochastic kinetic models are ubiquitous in physics, yet inferring their parameters from experimental data remains challenging. For determinis

Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology

Model ReleasesDGX agent

arXiv:2503.10629v2 Announce Type: replace Abstract: Adversarial attacks pose significant challenges for vision models in critical fields like healthcare, where reliability is essential. Although adver

Measuring What Matters: Synthetic Benchmarks for Concept Bottleneck Models

TutorialsDGX agent

arXiv:2606.04326v1 Announce Type: cross Abstract: Concept bottleneck models predict outcomes from high-level concepts detected in inputs. Although concepts provide a simple way to reap benefits from i

On TV Tokyo’s WBS (@wbs_tvtokyo) tonight I’ll be discussing Sakana AI’s upcoming 1T parameter model project, supported by METI’s GENIAC init…

Model ReleasesDGX agent

On TV Tokyo’s WBS (@wbs_tvtokyo) tonight I’ll be discussing Sakana AI’s upcoming 1T parameter model project, supported by METI’s GENIAC initiative. We are scaling up to build Japan’s first 1T paramete

OSCAR: Omni-Embodiment Skeleton-Conditioned World Action Model for Robotics

SafetyDGX agent

arXiv:2606.04463v1 Announce Type: new Abstract: We present OSCAR, a precise action-conditioned video world model that generalizes across different robot embodiments and enables robot policy evaluation

Overclocking Electrostatic Generative Models

ResearchDGX agent

arXiv:2509.22454v2 Announce Type: replace Abstract: Electrostatic generative models such as PFGM++ have recently emerged as a powerful framework, achieving competitive performance in image synthesis.

POLARIS: Guiding Small Models to Write Long Stories

SafetyDGX agent

arXiv:2606.04095v1 Announce Type: cross Abstract: Small open-weight models struggle at long-form creative writing: their generated stories either fall far short of the requested length, or their quali

PoliticsBench: Benchmarking Political Values in Large Language Models with Multi-Turn Roleplay

Model ReleasesDGX agent

arXiv:2603.23841v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) are increasingly used as primary sources of information, their potential for political bias may impact thei

Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning

Model ReleasesDGX agent

arXiv:2602.21103v2 Announce Type: replace Abstract: Advanced reasoning typically requires Chain-of-Thought prompting, which is accurate but incurs prohibitive latency and substantial test-time inferen

Robust-LLaVA: On the Effectiveness of Large-Scale Robust Image Encoders for Multi-modal Large Language Models

SafetyDGX agent

arXiv:2502.01576v2 Announce Type: replace Abstract: Multi-modal Large Language Models (MLLMs) excel in vision-language tasks but remain vulnerable to visual adversarial perturbations that can induce h

ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling

AgentsDGX agent

arXiv:2603.02697v2 Announce Type: replace-cross Abstract: This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works

Stateful Visual Encoders for Vision-Language Models

AgentsDGX agent

arXiv:2606.04433v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used in multi-image, multi-turn agentic settings where decisions depend on visual changes. However, in

Video2LoRA: Parametric Video Internalization for Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.04351v1 Announce Type: cross Abstract: Processing video in vision-language models is expensive: each frame occupies hundreds of tokens, and inference cost scales with every frame and every

We are excited to join Nvidia's Nemotron Coalition of leading AI labs working together to advance open frontier foundation models. To celebr…

Model ReleasesDGX agent

We are excited to join Nvidia's Nemotron Coalition of leading AI labs working together to advance open frontier foundation models. To celebrate we have partnered with @nvidia and @nebiustf to provide

What happened when one of our models found a counterexample to an 80-year-old Erdős conjecture? Researchers @alexwei_, @HongxunWu, and @wjmz…

Model ReleasesDGX agent

What happened when one of our models found a counterexample to an 80-year-old Erdős conjecture? Researchers @alexwei_, @HongxunWu, and @wjmzbmr1 shared the story on the OpenAI Podcast with @AndrewMayn

3 Jun 2026

AI Model Extraction Attacks: Bypassing Single-Client Assumptions in Defenses

ResearchDGX agent

arXiv:2606.03381v1 Announce Type: cross Abstract: Ensuring the protection of Artificial Intelligence (AI) models deployed in military Command and Control (C2) systems and critical infrastructure is es

Can Factual Opinions Be Edited (Manipulated) in Large Language Models?

Model ReleasesDGX agent

arXiv:2606.03096v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly integrated into various domains, making knowledge editing techniques crucial yet potentially hazardous. Cu

CREward: A Type-Specific Creativity Reward Model

Model ReleasesDGX agent

arXiv:2511.19995v2 Announce Type: replace Abstract: Creativity is a complex phenomenon. When it comes to representing and assessing creativity, treating it as a single undifferentiated quantity would

Cryo-Bench: Benchmarking Foundation Models for Cryosphere Applications

Model ReleasesDGX agent

arXiv:2603.01576v3 Announce Type: replace Abstract: Geo-Foundation Models (GFMs) have been evaluated across diverse Earth observation task including multiple domains and have demonstrated strong poten

dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive Caching

ResearchDGX agent

arXiv:2506.06295v2 Announce Type: replace-cross Abstract: Autoregressive Models (ARMs) have long dominated the landscape of Large Language Models. Recently, a new paradigm has emerged in the form of d

Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

Model ReleasesDGX agent

arXiv:2510.21011v3 Announce Type: replace-cross Abstract: As generative AI tools are increasingly used to portray people in professional roles, understanding their racial and gender representational b

Google's new Gemma 4 12B model is designed to run on any laptop with 16GB of RAM

Model ReleasesDGX agent

Google DeepMind released Gemma 4 12B, an open AI model that brings multimodal capabilities to everyday laptops by processing text, images, and audio natively without separate encoders. Small enough to

Make sure to update your runtime first! > lms runtime update --all Learn more about this model release https://x.com/googlegemma/status/2062…

Model ReleasesDGX agent

Make sure to update your runtime first! > lms runtime update --all Learn more about this model release https://x.com/googlegemma/status/2062202706882883696?s=20 Meet Gemma 4 12B! A unified, encoder-fr

← Previous
1…8485868788…1009
Next →