AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,860 results
10 Aug 2026

When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction

TutorialsDGX agent

arXiv:2505.16170v4 Announce Type: replace Abstract: We study the internal mechanisms that govern when LLMs choose to retract wrong answers, i.e., spontaneously and immediately acknowledge errors in th

Why Knowing Both Hops Is Not Enough: Understanding Two-Hop Generalization in Language Models

TutorialsDGX agent

arXiv:2608.07261v1 Announce Type: new Abstract: Large language models (LLMs) can solve complex multi-hop problems yet exhibit puzzling failures on simple two-hop queries: although a model may correctl

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

SafetyDGX agent

arXiv:2608.07267v1 Announce Type: new Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies t

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
7 Aug 2026

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

Model ReleasesDGX agent

arXiv:2608.05948v1 Announce Type: new Abstract: Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging as implicit s

Layer-wise Positional Bias in Short-Context Language Modeling

Model ReleasesDGX agent

arXiv:2601.04098v2 Announce Type: replace-cross Abstract: Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known as

OpenAI puts the brakes on a new model because it’s supposedly too powerful

Model ReleasesDGX agent

OpenAI says it is pausing 'internal activities' around an in-development AI model, Astra, because it doesn't yet meet new security standards the company is putting in place. The announcement follows i

6 Aug 2026

Above-ground Biomass Estimation with Geospatial Foundation Models

Model ReleasesDGX agent

arXiv:2608.04792v1 Announce Type: new Abstract: Accurate estimation of Above-Ground Biomass (AGB) from satellite imagery is essential for the large-scale monitoring of carbon stocks, yet it remains a

Enhancing Trustworthy Clinical Diagnosis Decision-Making in Large Language Models via Etiology-Aware Attention Supervision

Model ReleasesDGX agent

arXiv:2508.00285v2 Announce Type: replace Abstract: Objective: Large Language Models (LLMs) have demonstrated strong capabilities in medical text understanding and generation. However, their trustwort

Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth

Local AiDGX agent

arXiv:2608.05110v1 Announce Type: cross Abstract: Near-term quantum hardware limits circuit depth and often imposes geometrically local connectivity for quantum generative models, restricting the outp

Sources: Alibaba plans to ask heavy commercial users of its next Qwen open model for a share of revenue; Moonshot's Kimi K3 requires up to a 30% revenue share (Reuters)

Model ReleasesDGX agent

Reuters: Sources: Alibaba plans to ask heavy commercial users of its next Qwen open model for a share of revenue; Moonshot's Kimi K3 requires up to a 30% revenue share — Chinese technology giant Aliba

5 Aug 2026

AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction

Model ReleasesDGX agent

arXiv:2608.03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

ResearchDGX agent

arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m

Dynamically Allocating Evaluation Effort for Model Ranking

Model ReleasesDGX agent

arXiv:2608.03437v1 Announce Type: new Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing m

SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology

Local AiDGX agent

arXiv:2608.02803v1 Announce Type: cross Abstract: Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet its attentio

4 Aug 2026

Beyond VLAs: How World Action Models Reshape Robot Manipulation

SafetyDGX agent

World Action Models (WAMs) use video-based world modeling instead of vision‑language backbones, giving robots a learned physics engine that supports zero‑shot transfer to new tasks, environments, and

Company approved 128GB Mac for research proposal, best model?

Model ReleasesDGX agent

I‘m doing a research proposal at my company about running local LLMs to replace daily coding models. Qwen 3.6 27B (or 3.8 potentially) is widely seen as the best model in that 20-60GB space, is that s

Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark

Model ReleasesDGX agent

I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided to post a new one because of how well Deepseek did. I like the b

EEG-FM-Compass: Progress, Benchmarking, and Future Directions for EEG Foundation Models

Model ReleasesDGX agent

arXiv:2601.17883v3 Announce Type: replace-cross Abstract: Electroencephalography (EEG) foundation models (FMs) have recently emerged as a promising paradigm for brain-computer interfaces, aiming to le

Eigenvalues as a Metric for Memory Dynamics in Sequence Models

ResearchDGX agent

arXiv:2510.09379v2 Announce Type: replace Abstract: While softmax attention drives state-of-the-art performance in sequence modeling, its quadratic complexity motivates linear alternatives such as sta

Gaokerena: A Small Persian Medical Language Model Family

Model ReleasesDGX agent

arXiv:2608.00932v1 Announce Type: new Abstract: The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2608.02497v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models excel in robotic manipulation but suffer catastrophic performance drops when canonical instructions are simply parap

Large Causal Models for Temporal Causal Discovery

Model ReleasesDGX agent

arXiv:2602.18662v2 Announce Type: replace Abstract: Causal discovery for both cross-sectional and temporal data has traditionally followed a dataset-specific paradigm, where a new model is fitted for

Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models

Model ReleasesDGX agent

arXiv:2409.07163v3 Announce Type: replace-cross Abstract: Diffusion models have been widely employed in the field of 3D manipulation due to their efficient capability to learn distributions, allowing

MiniWorld: Democratizing the Training of Video World Models from Scratch

HardwareDGX agent

arXiv:2608.01127v1 Announce Type: new Abstract: Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through auto

Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models

Model ReleasesDGX agent

arXiv:2608.01624v1 Announce Type: new Abstract: Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable c

Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer

SafetyDGX agent

arXiv:2608.01585v1 Announce Type: new Abstract: Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source

SingLEM: Single-Channel Large EEG Model

Local AiDGX agent

arXiv:2509.17920v2 Announce Type: replace Abstract: Current deep learning models for electroencephalography (EEG) are often task-specific and depend on large labeled datasets, limiting their adaptabil

SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.00100v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly being evaluated for medical imaging, but many available benchmarks emphasize disease classification, repo

3 Aug 2026

All frontier AI models (including GPT-5.6 Sol, Fable 5, and Kimi K3) consistently fail to implement solvers for basic nonlinear PDEs correct…

Model ReleasesDGX agent

All frontier AI models (including GPT-5.6 Sol, Fable 5, and Kimi K3) consistently fail to implement solvers for basic nonlinear PDEs correctly, often introducing significant errors in numerical stabil

Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives

Model ReleasesDGX agent

arXiv:2607.29064v1 Announce Type: cross Abstract: Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclea

DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models

Model ReleasesDGX agent

arXiv:2607.28936v1 Announce Type: cross Abstract: Facial biometric identification relies on the distinctiveness of user attributes within a high-dimensional embedding space. However, the decision boun

Fisher Information, Training and Bias in Fourier Regression Models

SafetyDGX agent

arXiv:2510.06945v2 Announce Type: replace Abstract: Motivated by the growing interest in quantum machine learning, in particular quantum neural networks (QNNs), we study how recently introduced evalua

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

Model ReleasesDGX agent

arXiv:2607.17977v2 Announce Type: replace Abstract: We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and p

When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration

ResearchDGX agent

arXiv:2607.29240v1 Announce Type: cross Abstract: In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence of an atypi

2 Aug 2026

Another day and another full frontier model running on your computer. Been teething DeepSeek V4 Flash on over 60 employees at The Zero Human…

Model ReleasesDGX agent

Another day and another full frontier model running on your computer. Been teething DeepSeek V4 Flash on over 60 employees at The Zero Human Company for a few hours and it is stunning. Between Kimi K3

1 Aug 2026

DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities: ollama run d…

Model ReleasesDGX agent

DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities: ollama run deepseek-v4-flash:0731-cloud Use it with Claude Code: ollama

31 Jul 2026

Ask don't tell: Reducing sycophancy in large language models

SafetyDGX agent

arXiv:2602.23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an align

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

Model ReleasesDGX agent

arXiv:2607.28568v1 Announce Type: new Abstract: Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a

KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models

Model ReleasesDGX agent

arXiv:2607.28608v1 Announce Type: new Abstract: Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pip

MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models

Model ReleasesDGX agent

arXiv:2502.20780v2 Announce Type: replace-cross Abstract: The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which t

MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding

ResearchDGX agent

arXiv:2512.12307v5 Announce Type: replace Abstract: While deep learning methods have achieved impressive success in many vision benchmarks, it remains difficult to understand and explain the represent

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Model ReleasesDGX agent

arXiv:2607.28609v1 Announce Type: cross Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Veri

QuantWAMs: Calibrating at the Right Granularity for World Action Models

ResearchDGX agent

arXiv:2607.28405v1 Announce Type: cross Abstract: World Action Models (WAMs) jointly predict future observations and actions, but their iterative denoising and closed-loop execution make efficient dep

Scaling Vision-Language Models Is Not Enough to Mitigate Bias

Model ReleasesDGX agent

arXiv:2607.28211v1 Announce Type: new Abstract: Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly unde

Subtract or Replay? Exact Deletion from Language-Model Memory

Model ReleasesDGX agent

arXiv:2607.27539v1 Announce Type: cross Abstract: Exact deletion from persistent language-model memory depends on how that memory represents a record. Addressable influence can be removed by algebraic

The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models

SafetyDGX agent

arXiv:2602.08159v2 Announce Type: replace-cross Abstract: When a language model asserts that 'the capital of Australia is Sydney,' does it know this is wrong? Models assert misconceptions with the sam

TriShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement

Model ReleasesDGX agent

arXiv:2607.27940v1 Announce Type: cross Abstract: Federated fine-tuning of large language models (LLMs) enables collaborative training without exposing raw data. However, a recent attack, NeuroImprint

30 Jul 2026

Existence-Field Diffusion Model for Spatial Point Processes with Variable Cardinality

ResearchDGX agent

arXiv:2607.26428v1 Announce Type: new Abstract: We study generative modeling of spatial point processes (SPP), where both the number of points and their spatial configuration are governed by a joint d

Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark

Model ReleasesDGX agent

arXiv:2607.26993v1 Announce Type: new Abstract: Face presentation attack detection (PAD) remains challenging under cross-dataset evaluation, where domain shift degrades models trained on a single data

Learning Implicit Causal World Models from Multi-Agent Demonstrations

SafetyDGX agent

arXiv:2607.26336v1 Announce Type: new Abstract: In model-based reinforcement learning, world models exist as internal simulators, but their training often conflates statistical correlations with causa

P.A.I. — Sleek Native Desktop AI Overlayer for Local Ollama Models 🤖⚡

Model ReleasesDGX agent

Greetings Community! 👋 I hope everyone is doing well! I'm Tauhid — Senior EEE student from a Bangladeshi University Today I'd like to share an open-source project I’ve been developing called P.A.I. (P

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

SafetyDGX agent

arXiv:2607.26119v1 Announce Type: cross Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterpar

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

ResearchDGX agent

arXiv:2510.08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.

Walk through Paintings: Egocentric World Models from Internet Priors

ResearchDGX agent

arXiv:2601.15284v2 Announce Type: replace Abstract: What if a video generation model could not only imagine a plausible future, but the correct one -- accurately reflecting how the world changes with

29 Jul 2026

Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering

SafetyDGX agent

arXiv:2607.25479v1 Announce Type: cross Abstract: Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definitions, text

Building Large-Scale English-Romanian Literary Translation Resources with Open Models

Model ReleasesDGX agent

arXiv:2509.07829v4 Announce Type: replace-cross Abstract: Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small op

I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models

Local AiDGX agent

arXiv:2607.25522v1 Announce Type: cross Abstract: The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial progress has be

Influence of Prompt Engineering on Small Language Models for Guarded Query Routing

Model ReleasesDGX agent

arXiv:2607.24801v1 Announce Type: cross Abstract: We study the problem of guarded query routing, where we assume that a user query first meets a router that either determines the ideal endpoint for in

Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement

ResearchDGX agent

arXiv:2607.24765v1 Announce Type: cross Abstract: Large language models (LLMs) can give different answers to the same decision problem across runs, and reverse a decision when their own prior answer r

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

ResearchDGX agent

arXiv:2607.25948v1 Announce Type: cross Abstract: Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and vision-lang

← Previous
1…3435363738…998
Next →