AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,904 results
Model Releases

AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction

DGX agent

arXiv:2608.03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different

model-releasesarxiv-cs-ai
5 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

DGX agent

arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m

researcharxiv-cs-cl
5 Aug 2026
Model Releases

Dynamically Allocating Evaluation Effort for Model Ranking

DGX agent

arXiv:2608.03437v1 Announce Type: new Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing m

model-releasesarxiv-cs-cl
5 Aug 2026
Local Ai

SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology

DGX agent

arXiv:2608.02803v1 Announce Type: cross Abstract: Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet its attentio

local-aiarxiv-cs-ai
5 Aug 2026
Safety

Beyond VLAs: How World Action Models Reshape Robot Manipulation

DGX agent

World Action Models (WAMs) use video-based world modeling instead of vision‑language backbones, giving robots a learned physics engine that supports zero‑shot transfer to new tasks, environments, and

safetynvidia-developer
4 Aug 2026
Model Releases

Company approved 128GB Mac for research proposal, best model?

DGX agent

I‘m doing a research proposal at my company about running local LLMs to replace daily coding models. Qwen 3.6 27B (or 3.8 potentially) is widely seen as the best model in that 20-60GB space, is that s

model-releasesr-localllama
4 Aug 2026
Model Releases

Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark

DGX agent

I really like to use this one SQL benchmark when testing new models. I had another post some time ago with my benchmarks, but I decided to post a new one because of how well Deepseek did. I like the b

model-releasesr-localllama
4 Aug 2026
Model Releases

EEG-FM-Compass: Progress, Benchmarking, and Future Directions for EEG Foundation Models

DGX agent

arXiv:2601.17883v3 Announce Type: replace-cross Abstract: Electroencephalography (EEG) foundation models (FMs) have recently emerged as a promising paradigm for brain-computer interfaces, aiming to le

model-releasesarxiv-cs-cv
4 Aug 2026
Research

Eigenvalues as a Metric for Memory Dynamics in Sequence Models

DGX agent

arXiv:2510.09379v2 Announce Type: replace Abstract: While softmax attention drives state-of-the-art performance in sequence modeling, its quadratic complexity motivates linear alternatives such as sta

researcharxiv-cs-lg
4 Aug 2026
Model Releases

Gaokerena: A Small Persian Medical Language Model Family

DGX agent

arXiv:2608.00932v1 Announce Type: new Abstract: The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused

model-releasesarxiv-cs-cl
4 Aug 2026
Model Releases

Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models

DGX agent

arXiv:2608.02497v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models excel in robotic manipulation but suffer catastrophic performance drops when canonical instructions are simply parap

model-releasesarxiv-cs-ro
4 Aug 2026
Model Releases

Large Causal Models for Temporal Causal Discovery

DGX agent

arXiv:2602.18662v2 Announce Type: replace Abstract: Causal discovery for both cross-sectional and temporal data has traditionally followed a dataset-specific paradigm, where a new model is fitted for

model-releasesarxiv-cs-lg
4 Aug 2026
Model Releases

Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State Models

DGX agent

arXiv:2409.07163v3 Announce Type: replace-cross Abstract: Diffusion models have been widely employed in the field of 3D manipulation due to their efficient capability to learn distributions, allowing

model-releasesarxiv-cs-cv
4 Aug 2026
Hardware

MiniWorld: Democratizing the Training of Video World Models from Scratch

DGX agent

arXiv:2608.01127v1 Announce Type: new Abstract: Video world models predict future observations conditioned on historical observations and control signals, enabling long-horizon generation through auto

hardwarearxiv-cs-cv
4 Aug 2026
Model Releases

Not the Dimension, the Norm: What Matters in Gradient-Free Weight Perturbation of Language Models

DGX agent

arXiv:2608.01624v1 Announce Type: new Abstract: Adapting a language model to a task no longer requires training all of its weights, and a line of parameter-efficient methods has driven the trainable c

model-releasesarxiv-cs-cl
4 Aug 2026
Safety

Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer

DGX agent

arXiv:2608.01585v1 Announce Type: new Abstract: Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source

safetyarxiv-cs-cl
4 Aug 2026
Local Ai

SingLEM: Single-Channel Large EEG Model

DGX agent

arXiv:2509.17920v2 Announce Type: replace Abstract: Current deep learning models for electroencephalography (EEG) are often task-specific and depend on large labeled datasets, limiting their adaptabil

local-aiarxiv-cs-lg
4 Aug 2026
Model Releases

SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models

DGX agent

arXiv:2608.00100v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly being evaluated for medical imaging, but many available benchmarks emphasize disease classification, repo

model-releasesarxiv-cs-cv
4 Aug 2026
Model Releases

All frontier AI models (including GPT-5.6 Sol, Fable 5, and Kimi K3) consistently fail to implement solvers for basic nonlinear PDEs correct…

DGX agent

All frontier AI models (including GPT-5.6 Sol, Fable 5, and Kimi K3) consistently fail to implement solvers for basic nonlinear PDEs correctly, often introducing significant errors in numerical stabil

model-releasesgary-marcus--x
3 Aug 2026
Model Releases

Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives

DGX agent

arXiv:2607.29064v1 Announce Type: cross Abstract: Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclea

model-releasesarxiv-cs-ai
3 Aug 2026
Model Releases

DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models

DGX agent

arXiv:2607.28936v1 Announce Type: cross Abstract: Facial biometric identification relies on the distinctiveness of user attributes within a high-dimensional embedding space. However, the decision boun

model-releasesarxiv-cs-ai
3 Aug 2026
Safety

Fisher Information, Training and Bias in Fourier Regression Models

DGX agent

arXiv:2510.06945v2 Announce Type: replace Abstract: Motivated by the growing interest in quantum machine learning, in particular quantum neural networks (QNNs), we study how recently introduced evalua

safetyarxiv-cs-lg
3 Aug 2026
Model Releases

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

DGX agent

arXiv:2607.17977v2 Announce Type: replace Abstract: We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and p

model-releasesarxiv-cs-ro
3 Aug 2026
Research

When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration

DGX agent

arXiv:2607.29240v1 Announce Type: cross Abstract: In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence of an atypi

researcharxiv-cs-ai
3 Aug 2026
Model Releases

Another day and another full frontier model running on your computer. Been teething DeepSeek V4 Flash on over 60 employees at The Zero Human…

DGX agent

Another day and another full frontier model running on your computer. Been teething DeepSeek V4 Flash on over 60 employees at The Zero Human Company for a few hours and it is stunning. Between Kimi K3

model-releasesclem-delangue--x
2 Aug 2026
Model Releases

DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities: ollama run d…

DGX agent

DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities: ollama run deepseek-v4-flash:0731-cloud Use it with Claude Code: ollama

model-releasesollama--x
1 Aug 2026
Safety

Ask don't tell: Reducing sycophancy in large language models

DGX agent

arXiv:2602.23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an align

safetyarxiv-cs-ai
31 Jul 2026
Model Releases

Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

DGX agent

arXiv:2607.28568v1 Announce Type: new Abstract: Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a

model-releasesarxiv-cs-cl
31 Jul 2026
Model Releases

KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models

DGX agent

arXiv:2607.28608v1 Announce Type: new Abstract: Clinical risk models routinely achieve strong aggregate performance while producing materially different error rates across patient subgroups. Audit pip

model-releasesarxiv-cs-lg
31 Jul 2026
Model Releases

MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models

DGX agent

arXiv:2502.20780v2 Announce Type: replace-cross Abstract: The increasing use of vision-language models (VLMs) in healthcare applications presents great challenges related to hallucinations, in which t

model-releasesarxiv-cs-cl
31 Jul 2026
Research

MRD: Using Physically Based Differentiable Rendering to Probe Vision Models for 3D Scene Understanding

DGX agent

arXiv:2512.12307v5 Announce Type: replace Abstract: While deep learning methods have achieved impressive success in many vision benchmarks, it remains difficult to understand and explain the represent

researcharxiv-cs-cv
31 Jul 2026
Model Releases

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

DGX agent

arXiv:2607.28609v1 Announce Type: cross Abstract: Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Veri

model-releasesarxiv-cs-cl
31 Jul 2026
Research

QuantWAMs: Calibrating at the Right Granularity for World Action Models

DGX agent

arXiv:2607.28405v1 Announce Type: cross Abstract: World Action Models (WAMs) jointly predict future observations and actions, but their iterative denoising and closed-loop execution make efficient dep

researcharxiv-cs-lg
31 Jul 2026
Model Releases

Scaling Vision-Language Models Is Not Enough to Mitigate Bias

DGX agent

arXiv:2607.28211v1 Announce Type: new Abstract: Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly unde

model-releasesarxiv-cs-cv
31 Jul 2026
Model Releases

Subtract or Replay? Exact Deletion from Language-Model Memory

DGX agent

arXiv:2607.27539v1 Announce Type: cross Abstract: Exact deletion from persistent language-model memory depends on how that memory represents a record. Addressable influence can be removed by algebraic

model-releasesarxiv-cs-cl
31 Jul 2026
Safety

The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models

DGX agent

arXiv:2602.08159v2 Announce Type: replace-cross Abstract: When a language model asserts that 'the capital of Australia is Sydney,' does it know this is wrong? Models assert misconceptions with the sam

safetyarxiv-cs-cl
31 Jul 2026
Model Releases

TriShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement

DGX agent

arXiv:2607.27940v1 Announce Type: cross Abstract: Federated fine-tuning of large language models (LLMs) enables collaborative training without exposing raw data. However, a recent attack, NeuroImprint

model-releasesarxiv-cs-cl
31 Jul 2026
Research

Existence-Field Diffusion Model for Spatial Point Processes with Variable Cardinality

DGX agent

arXiv:2607.26428v1 Announce Type: new Abstract: We study generative modeling of spatial point processes (SPP), where both the number of points and their spatial configuration are governed by a joint d

researcharxiv-cs-lg
30 Jul 2026
Model Releases

Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark

DGX agent

arXiv:2607.26993v1 Announce Type: new Abstract: Face presentation attack detection (PAD) remains challenging under cross-dataset evaluation, where domain shift degrades models trained on a single data

model-releasesarxiv-cs-lg
30 Jul 2026
Safety

Learning Implicit Causal World Models from Multi-Agent Demonstrations

DGX agent

arXiv:2607.26336v1 Announce Type: new Abstract: In model-based reinforcement learning, world models exist as internal simulators, but their training often conflates statistical correlations with causa

safetyarxiv-cs-lg
30 Jul 2026
Model Releases

P.A.I. — Sleek Native Desktop AI Overlayer for Local Ollama Models 🤖⚡

DGX agent

Greetings Community! 👋 I hope everyone is doing well! I'm Tauhid — Senior EEE student from a Bangladeshi University Today I'd like to share an open-source project I’ve been developing called P.A.I. (P

model-releasesr-ollama
30 Jul 2026
Safety

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

DGX agent

arXiv:2607.26119v1 Announce Type: cross Abstract: Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterpar

safetyarxiv-cs-cl
30 Jul 2026
Research

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

DGX agent

arXiv:2510.08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.

researcharxiv-cs-cl
30 Jul 2026
Research

Walk through Paintings: Egocentric World Models from Internet Priors

DGX agent

arXiv:2601.15284v2 Announce Type: replace Abstract: What if a video generation model could not only imagine a plausible future, but the correct one -- accurately reflecting how the world changes with

researcharxiv-cs-cv
30 Jul 2026
Safety

Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering

DGX agent

arXiv:2607.25479v1 Announce Type: cross Abstract: Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definitions, text

safetyarxiv-cs-ai
29 Jul 2026
Model Releases

Building Large-Scale English-Romanian Literary Translation Resources with Open Models

DGX agent

arXiv:2509.07829v4 Announce Type: replace-cross Abstract: Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small op

model-releasesarxiv-cs-ai
29 Jul 2026
Local Ai

I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models

DGX agent

arXiv:2607.25522v1 Announce Type: cross Abstract: The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial progress has be

local-aiarxiv-cs-ai
29 Jul 2026
Model Releases

Influence of Prompt Engineering on Small Language Models for Guarded Query Routing

DGX agent

arXiv:2607.24801v1 Announce Type: cross Abstract: We study the problem of guarded query routing, where we assume that a user query first meets a router that either determines the ideal endpoint for in

model-releasesarxiv-cs-cl
29 Jul 2026
← Previous
1…4344454647…1248
Next →