AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,588
  • Agents7,266
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,577
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,588Total entries
1Added by human
84,587Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,585 results
Model Releases

Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting

DGX agent

arXiv:2607.02467v1 Announce Type: cross Abstract: Whether pairing people with AI helps or hurts is usually reported as a single average effect. Using a real-money prediction market (Polymarket) as an

model-releasesarxiv-cs-ai
3 Jul 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

I kept asking Claude Fable to make the game 'more AAA' over and over again. The results are... interesting. In Claude's view, this meant upg…

DGX agent

I kept asking Claude Fable to make the game 'more AAA' over and over again. The results are... interesting. In Claude's view, this meant upgrading graphics, boss fights, mechanics adding custom sounds

model-releasesethan-mollick--x
3 Jul 2026
Model Releases

InduceKV: Fixed-Footprint Continual Adaptation of Multimodal LLMs via Inducing KV Memories

DGX agent

arXiv:2607.02010v1 Announce Type: new Abstract: Multimodal large language models must adapt to evolving tasks and domains, yet continual improvement under bounded deployment footprint remains difficul

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Influence of Radial Basis Activation Functions on Intelligent Controller for Robotic Manipulators

DGX agent

arXiv:2607.02167v1 Announce Type: cross Abstract: This paper presents an intelligent control framework for trajectory tracking of robotic manipulators using radial basis function (RBF) neural networks

model-releasesarxiv-cs-ro
3 Jul 2026
Model Releases

IonSense-QKG: A Quantum-Readiness Metadata Framework for Lithium-Ion Battery Dataset Discovery

DGX agent

arXiv:2607.01286v1 Announce Type: new Abstract: Public lithium-ion battery datasets are increasingly used for state-of-health estimation, remaining-useful-life prediction, anomaly detection, electroch

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs

DGX agent

arXiv:2607.01431v1 Announce Type: cross Abstract: We introduce ISOSCI, a benchmark of isomorphic cross-domain science problem pairs that separates reasoning ability from domain knowledge retrieval in

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

June 2026 newsletter

DGX agent

The June edition of my sponsors-only monthly newsletter is out. If you are a sponsor (or if you start a sponsorship now) you can access it here. This month: Claude Fable 5, GPT-5.6, and US export rest

model-releasessimon-willison
3 Jul 2026
Model Releases

LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning

DGX agent

arXiv:2607.02513v1 Announce Type: cross Abstract: LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal met

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models

DGX agent

arXiv:2504.02327v2 Announce Type: replace Abstract: Natural Language to SQL (NL2SQL) aims to translate natural language queries into executable SQL statements, offering non-expert users intuitive acce

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

DGX agent

arXiv:2607.02466v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, instructions,

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens

DGX agent

arXiv:2507.02964v2 Announce Type: replace-cross Abstract: The increasing scale of AI workloads demands High-Performance Computing (HPC) infrastructure and training methodologies that are both scalable

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Liquid Latent State Dynamics for Interpretable Turbofan Degradation Modeling

DGX agent

arXiv:2607.01986v1 Announce Type: new Abstract: Multivariate time-series models for prognostics are often evaluated by point prediction accuracy, yet their internal states rarely expose a coherent deg

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability

DGX agent

arXiv:2607.01247v1 Announce Type: cross Abstract: Open-ended mathematics exams are valuable because they assess reasoning, proof construction, algorithmic thinking, and communication of intermediate s

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Locality-Aware Continual Unlearning for Diffusion Models

DGX agent

arXiv:2512.02657v2 Announce Type: replace-cross Abstract: Real-world deployment of text-to-image diffusion models requires continual concept removal as new privacy, copyright, or safety obligations ar

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction

DGX agent

arXiv:2607.01764v1 Announce Type: new Abstract: Repository-level vulnerability reproduction is a demanding software engineering (SE) task: an agent must inspect a codebase, infer the input grammar tha

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation

DGX agent

arXiv:2508.16674v2 Announce Type: replace-cross Abstract: Medical report understanding from real-world document images is essential for generating patient-facing explanations and enabling structured i

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding

DGX agent

arXiv:2607.01751v1 Announce Type: cross Abstract: Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at the right ti

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Meta-Benchmarks for Financial-Services LLM Evaluation

DGX agent

arXiv:2607.01740v1 Announce Type: new Abstract: Public LLM leaderboards optimise for global average performance and do not capture the specific cognitive demands of financial-services work: a model th

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Meta to release new AI model with advanced coding capabilities ‘soon’

DGX agent

Meta Platforms Inc. is gearing up to release a new version of its flagship Muse Spark artificial intelligence model. Alexandr Wang, the company’s chief AI officer, wrote on X today that the update wil

model-releasessiliconangle
3 Jul 2026
Model Releases

MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2506.09105v3 Announce Type: replace-cross Abstract: We present MetaTT, a Tensor Train (TT) adapter framework for fine-tuning of pre-trained transformers. MetaTT enables flexible and parameter-ef

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models

DGX agent

arXiv:2607.01844v1 Announce Type: cross Abstract: This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and specializes va

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MKGR: Multimodal Knowledge-Graph Representation Learning for Cold-Start Protein-Protein Interaction Prediction

DGX agent

arXiv:2607.01627v1 Announce Type: cross Abstract: Accurate protein-protein interaction (PPI) prediction is central to functional genomics, disease mechanism discovery, and drug development. A difficul

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

DGX agent

arXiv:2607.01813v1 Announce Type: cross Abstract: Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support

DGX agent

arXiv:2607.01814v1 Announce Type: new Abstract: Traditional Chinese Medicine (TCM) diagnosis, particularly through tongue inspection, faces persistent challenges in subjectivity and reproducibility. T

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Model Merging as Probabilistic Inference in Fine-Tuning Parameter Space

DGX agent

arXiv:2607.01689v1 Announce Type: cross Abstract: Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.~Most existing appr

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

DGX agent

arXiv:2607.01420v1 Announce Type: cross Abstract: As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for user trust and

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation

DGX agent

arXiv:2604.04532v2 Announce Type: replace-cross Abstract: Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's language

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

mupscaling small models: Principled warm starts and hyperparameter transfer

DGX agent

arXiv:2602.10545v2 Announce Type: replace-cross Abstract: Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve effic

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding

DGX agent

arXiv:2601.01095v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand tempo

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Neuro-Symbolic Safety Guidance for Vision-Language-Action Models via Constrained Flow Matching

DGX agent

arXiv:2607.01378v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated promising generalization capabilities across robotic manipulation tasks, yet their real-world depl

model-releasesarxiv-cs-ro
3 Jul 2026
Model Releases

NEUROSYMLAND: Neuro-Symbolic Landing-Site Assessment for Robust and Edge-Deployable UAV Autonomy

DGX agent

arXiv:2607.02277v1 Announce Type: new Abstract: Safe landing-site assessment in unstructured environments remains a key challenge for autonomous UAV deployment, as vision-only learning approaches ofte

model-releasesarxiv-cs-ro
3 Jul 2026
Model Releases

Office Comprehension Benchmark

DGX agent

arXiv:2607.01245v1 Announce Type: cross Abstract: We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint comprehension

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

OmniGAIA: Towards Native Omni-Modal AI Agents

DGX agent

arXiv:2602.22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to i

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

On the Limits of Steering Vectors for Preference-Aligned Generation

DGX agent

arXiv:2607.01802v1 Announce Type: new Abstract: Steering vectors have emerged as a promising approach to controlled text generation, offering interpretable, training-free mechanisms for shaping model

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain

DGX agent

arXiv:2607.01444v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models offer inference speedups via selective activation but impose substantial memory requirements because the whole network

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

One More Time: Revisiting Neural Quantum States from a Reinforcement Learning Perspective

DGX agent

arXiv:2607.02292v1 Announce Type: new Abstract: Neural quantum states (NQS) provide a flexible and scalable framework for approximating quantum many-body wavefunctions. Among NQS parameterizations, au

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Open Source AI Gap Map

DGX agent

Open Source AI Gap Map Current AI is 'a global partnership building a public option for AI', founded as a non-profit at the AI Action Summit in Paris in February 2025 and backed by serious capital ($4

model-releasessimon-willison
3 Jul 2026
Model Releases

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets

DGX agent

arXiv:2607.02047v1 Announce Type: cross Abstract: Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolated prompts.

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration

DGX agent

arXiv:2607.01531v1 Announce Type: new Abstract: Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networ

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers

DGX agent

arXiv:2607.02461v1 Announce Type: cross Abstract: Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make infe

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

PACE: A Proxy for Agentic Capability Evaluation

DGX agent

arXiv:2607.02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation c

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation

DGX agent

arXiv:2607.01883v1 Announce Type: new Abstract: Code is the medium through which large language models generate structured artifacts: charts, scientific figures, vector graphics, CAD models, 3D scenes

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Parameter Golf: What Really Works?

DGX agent

arXiv:2607.01517v1 Announce Type: new Abstract: How far can a language model improve under a strict artifact budget? Parameter Golf posed this question as an open community challenge in which particip

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

DGX agent

arXiv:2607.01754v1 Announce Type: new Abstract: On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribu

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Phonikud: Overcoming Phonetic Underspecification for Hebrew Text-To-Speech

DGX agent

arXiv:2506.12311v4 Announce Type: replace Abstract: Text-to-speech (TTS) for Modern Hebrew is challenged by the language's orthographic complexity, with existing solutions ignoring underspecified phon

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

DGX agent

arXiv:2607.01938v1 Announce Type: cross Abstract: Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-language-action

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Population-Scale Segmentation of Penile Tissue in DIXON MRI using Deep Learning for Quantitative Phenotyping in Male Reproductive Health

DGX agent

arXiv:2607.02127v1 Announce Type: cross Abstract: Penile measurement is clinically relevant across male reproductive and urogenital health, including conditions such as micropenis, congenital and endo

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Power Systems Agent Benchmark: Executable Evaluation of AI Agents in Electric Power Engineering

DGX agent

arXiv:2606.20950v2 Announce Type: replace Abstract: Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a prominent way

model-releasesarxiv-cs-ai
3 Jul 2026
← Previous
1…134135136137138…471
Next →