AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,929 results
25 Jun 2026

Certification of Machine Learning Models via Directional Sharpness

ResearchDGX agent

arXiv:2606.25004v1 Announce Type: new Abstract: In machine learning, model certification has been identified as an important method for gaining assurance about a model's trustworthiness and quality. A

Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models

Model ReleasesDGX agent

arXiv:2606.25324v1 Announce Type: new Abstract: The computational complexity of Transformers scales quadratically with the number of tokens, which significantly constrains the efficiency of vision mod

EPTS: Elastic Post-Training Sparsity for Efficient Large Language Model Compression

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.25285v1 Announce Type: new Abstract: Post-Training Sparsity (PTS) has emerged as a crucial paradigm for compressing Large Language Models to facilitate efficient deployment on resource-cons

Perfect Detection, Failed Control: The Geometry of Knowing vs. Steering in Language Models

Model ReleasesDGX agent

arXiv:2606.24952v1 Announce Type: new Abstract: A central aspiration of mechanistic interpretability is controllability: if we know where a behavior is represented in a model's activations, we should

Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models

AgentsDGX agent

arXiv:2606.25519v1 Announce Type: cross Abstract: Quantization is widely used to reduce the inference cost of large language models, but its effect on reasoning models is not fully captured by final-a

Robustness assessment of large audio language models in multiple-choice evaluation

Model ReleasesDGX agent

arXiv:2510.04584v2 Announce Type: replace Abstract: Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. How

RoDyn: Taming Interactive Robot-Dynamic 2.5D World Model for Robotic Manipulation

AgentsDGX agent

arXiv:2510.09036v2 Announce Type: replace Abstract: Learned world models hold significant potential as neural simulators for robotic manipulation. However, prevalent 2D video-based models inherently l

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2606.26079v1 Announce Type: new Abstract: Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling c

Small Initialization Matters for Large Language Models

Model ReleasesDGX agent

arXiv:2606.17945v2 Announce Type: replace Abstract: Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered. Although p

Towards Scalable Multi-Task Reinforcement Learning with Large Decision Models

SafetyDGX agent

arXiv:2606.24962v1 Announce Type: new Abstract: Recent progress in large-scale sequence modeling has shown that a single model can learn useful representations across highly diverse data distributions

24 Jun 2026

AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability

Model ReleasesDGX agent

arXiv:2606.24589v1 Announce Type: new Abstract: Scaling adversarial evaluation of large language models requires both a method for generating hard inputs and a reliable way to confirm that resulting f

Business as Rulesual: A Benchmark and Framework for Business Rule Flow Modeling with LLMs

Model ReleasesDGX agent

arXiv:2505.18542v4 Announce Type: replace Abstract: Extracting structured procedural knowledge from unstructured business documents is a critical yet unresolved bottleneck in process automation. While

Can Scale Save Us From Plasticity Loss in Large Language Models?

Model ReleasesDGX agent

arXiv:2606.24752v1 Announce Type: new Abstract: The loss of plasticity - the ability of a network to learn new information after having already learned older information - is a fundamental challenge i

Extended pseudo-spectral physics-informed neural networks for phase-field models

Model ReleasesDGX agent

arXiv:2606.24660v1 Announce Type: cross Abstract: Phase-field models play a central role in the continuum description of phase separation, in which the bulk free-energy density and the interfacial thi

HyMaTE: A Hybrid Mamba and Transformer Model for EHR Representation Learning

ApplicationsDGX agent

arXiv:2509.24118v2 Announce Type: replace Abstract: Electronic health Records (EHRs) have become a cornerstone in modern-day healthcare. They are a crucial part for analyzing the progression of patien

Revealing Training Data Exposure in Vision Language Large Models via Parameter Gradients

Model ReleasesDGX agent

arXiv:2606.24774v1 Announce Type: new Abstract: Vision-Language Large Models (VLLMs) trained on massive crawled corpora raise pressing copyright and data-provenance concerns. These concerns are partic

Sentence-Level Contextual Entrainment in Large Language Models

ResearchDGX agent

arXiv:2606.24077v1 Announce Type: new Abstract: Contextual entrainment, which is a newly discovered phenomenon in large language models (LLMs), refers to the tendency of a model to assign higher proba

The next scaling law is multi-agent swarms. Mixture of models.

AgentsDGX agent

The next scaling law is multi-agent swarms. Mixture of models. Introducing Sakana Fugu: A full multi-agent orchestration system accessible via a single model API. Our ‘Fugu Ultra’ model matches the pe

Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical Assessment

Model ReleasesDGX agent

arXiv:2606.24331v1 Announce Type: new Abstract: Transformer-based language models have become the default substrate for natural language processing and the pace of new releases has made it hard for pr

23 Jun 2026

A-Evolve-Training: Autonomous Post-Training of a 30B Model

Model ReleasesDGX agent

arXiv:2606.20657v1 Announce Type: cross Abstract: Post-training a frontier model is normally weeks of human work: proposing data and recipe changes, launching runs, reading evals, deciding what to kee

A Linear Fractional Transformation Model and Calibration Method for Light Field Camera

Model ReleasesDGX agent

arXiv:2511.03962v2 Announce Type: replace Abstract: Accurate intrinsic calibration is a crucial yet challenging prerequisite for 3D reconstruction using light field cameras. Existing calibration model

Attacking the Trusted Imagination: Oracle-Level Integrity Attacks on Imagine-then-Act World Models

Model ReleasesDGX agent

arXiv:2606.22966v1 Announce Type: new Abstract: Many recent vision-language-action (VLA) policies adopt an imagine-then-act design. A world-action model (WAM) first imagines a short future as a latent

Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models

SafetyDGX agent

arXiv:2606.21645v1 Announce Type: cross Abstract: Large language models (LLMs) can readily reproduce conventional expressions, yet their ability to model gradient frequency distributions remains under

Diffusion-Driven State Space Models

TutorialsDGX agent

arXiv:2606.21036v1 Announce Type: cross Abstract: In many domains, practitioners seek models that produce accurate forecasts while faithfully capturing latent system dynamics. Existing approaches typi

Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability

Model ReleasesDGX agent

arXiv:2606.22269v1 Announce Type: cross Abstract: We investigate the translation quality of current large language models (LLMs) for English-to-Hausa and English-to-Fongbe - two typologically distinct

How Well Do Self-Supervised Speech Models Encode Age and Gender in Children's Speech? A Layer-Wise Analysis Across Multiple Architectures

Model ReleasesDGX agent

arXiv:2606.22177v1 Announce Type: cross Abstract: Self-supervised learning (SSL) models have become a central component of modern speech processing systems, as they enable the learning of rich acousti

L20-Edu-135M: An Auditable Single-GPU Study of Data-Efficient Small Language Modeling

Model ReleasesDGX agent

arXiv:2606.22189v1 Announce Type: new Abstract: Small language models are cheap to serve and feasible on local hardware, but strong public 135M-class systems are commonly trained with hundreds of bill

Learning-Based Modeling of Soft Robots via Cosserat Rod Theory

ResearchDGX agent

arXiv:2606.20958v1 Announce Type: new Abstract: Modeling soft robot dynamics is challenging due to their continuum structure and typically nonlinear dynamics. Creating models based on first-order prin

LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.23686v1 Announce Type: new Abstract: Despite the impressive manipulation capabilities of Vision-Language-Action (VLA) models, their operational safety under strict constraints remains large

Load Testing for Machine Learning Model Serving Systems at Scale

Model ReleasesDGX agent

arXiv:2606.22013v1 Announce Type: new Abstract: Machine learning (ML) model serving has become a dominant consumer of GPU infrastructure, yet capacity planning in these systems remains largely ad hoc.

LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2509.23729v3 Announce Type: replace Abstract: Large Language Models (LLMs) with multimodal capabilities have revolutionized vision-language tasks, but their deployment often requires huge memory

Model Inversion meets Cryptographic Fuzzy Extractors

ResearchDGX agent

arXiv:2510.25687v3 Announce Type: replace-cross Abstract: Model inversion attacks pose an open challenge to privacy-sensitive applications that use machine learning (ML) models. For example, face auth

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation

SafetyDGX agent

arXiv:2606.18375v2 Announce Type: replace Abstract: World foundation models (WFMs) are powerful simulators, yet they predominantly operate in a single-view setting and lack the multi-view 3D consisten

Post-Training Speech Enhancement Language Models with Perceptual Rewards

Model ReleasesDGX agent

arXiv:2606.21458v1 Announce Type: new Abstract: Speech enhancement language models achieve strong results when trained on discrete audio tokens, but their optimization relies on token-level cross-entr

Residue-Level Attributions in Protein Language Models Do Not Recover Allergen Epitopes

Model ReleasesDGX agent

arXiv:2606.22181v1 Announce Type: new Abstract: Deep allergenicity classifiers are increasingly used in safety screening of novel foods, and recent protein language models have substantially improved

Translating Inference-Time Control to Radiology Vision-Language Models: Activation Steering for Pneumonia Classification on Chest X-rays

Model ReleasesDGX agent

arXiv:2606.20852v1 Announce Type: new Abstract: Inference-time engineering can alter model behavior without fine-tuning. However, its utility for improving diagnostic performance in medical vision-lan

VRPO: Rethinking Value Modeling for Robust RL under Noisy Supervision in LLM Post-Training

SafetyDGX agent

arXiv:2508.03058v2 Announce Type: replace Abstract: Reinforcement Learning (RL) in real-world environments often suffers from ambiguous or incomplete reward supervision, which undermines policy stabil

22 Jun 2026

Fugu stands shoulder-to-shoulder with leading models like Fable and Mythos across the industry's most rigorous engineering, scientific, and …

ApplicationsDGX agent

Fugu stands shoulder-to-shoulder with leading models like Fable and Mythos across the industry's most rigorous engineering, scientific, and reasoning benchmarks. Read the full blog: https://sakana.ai/

21 Jun 2026

Yes, is model agnostic and general purpose

Model ReleasesDGX agent

Harrison Chase confirms that 'Yes' (likely referring to LangChain or a related tool/framework) is model-agnostic and designed as a general-purpose solution, meaning it can work across different AI mod

20 Jun 2026

Hot take: GLM 5.2 might be the first open/public model that actually changes the enterprise AI cost equation. I played with it for a few hou…

HardwareDGX agent

Hot take: GLM 5.2 might be the first open/public model that actually changes the enterprise AI cost equation. I played with it for a few hours today after friends told me to stop ignoring it. I expect

11 Jun 2026

Google open-sources speedy DiffusionGemma text diffusion model

Model ReleasesDGX agent

Google LLC today released DiffusionGemma, a large language model based on an emerging machine learning approach known as text diffusion. The company says the algorithm can generate text four times fas

OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.12169v1 Announce Type: cross Abstract: High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not ju

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task

AgentsDGX agent

arXiv:2606.11830v1 Announce Type: new Abstract: Background. Large language models and AI agents are increasingly used to support biomedical research, but native model outputs may omit key analytical s

TimeRouter: Efficient and Adaptive Routing of Time-Series Foundation Models

Model ReleasesDGX agent

arXiv:2606.11625v1 Announce Type: new Abstract: Time-series foundation models (TSFMs) are increasingly explored as predictive experts within emerging agentic time-series systems. However, TSFMs exhibi

When Does Language Matter? Multilingual Instructions Reveal Step-wise Language Sensitivity in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.11906v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong performance in language-conditioned robotic manipulation, yet their robustness to linguistic varia

World Pilot: Steering Vision-Language-Action Models with World-Action Priors

Model ReleasesDGX agent

arXiv:2606.12403v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models inherit semantic grounding from large-scale pretraining and perform competently across in-distribution manipulation

10 Jun 2026

Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

Model ReleasesDGX agent

arXiv:2606.10040v1 Announce Type: new Abstract: World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. Howeve

Entropy, Disagreement, and the Limits of Foundation Models in Genomics

TutorialsDGX agent

arXiv:2604.04287v2 Announce Type: replace-cross Abstract: Foundation models in genomics have shown mixed success compared to their counterparts in natural language processing. Yet, the reasons for the

FADA: Accessible fetal ultrasound interpretation and annotation with a selectively distilled unified vision-language model

Model ReleasesDGX agent

arXiv:2606.11106v1 Announce Type: cross Abstract: A global shortage of trained sonographers limits prenatal ultrasound screening in low- and middle-income countries, where over half of pregnant women

Improving Adversarial Transferability on Vision-Language Pre-training Models via Surrogate-Specific Bias Correction

SafetyDGX agent

arXiv:2606.10571v1 Announce Type: cross Abstract: Adversarial examples reveal vulnerabilities in Vision-Language Pre-training (VLP) models and provide insights for improving robustness. A key property

P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning

Model ReleasesDGX agent

arXiv:2606.11152v1 Announce Type: new Abstract: Multimodal large language models can write code to produce complex programs as well as use programs to do 3D modeling, which opens up a new avenue for 3

Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models

Model ReleasesDGX agent

arXiv:2606.10949v1 Announce Type: new Abstract: Persistent memory systems promise to make LLMs more helpful by storing user beliefs over time. We show they also make models less correct by systematica

SARM2: Multi-Task Stage Aware Reward Modeling for Self Improving Robotic Manipulation

Model ReleasesDGX agent

arXiv:2606.10305v1 Announce Type: new Abstract: Fine-tuning vision-language-action (VLA) policies for long-horizon manipulation still relies heavily on behavior cloning, which requires costly high-qua

The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models

Model ReleasesDGX agent

arXiv:2606.11082v1 Announce Type: new Abstract: This study investigates cross-lingual distributional skew (the Shibboleth Effect) in frontier large language models (LLMs) subjected to sustained advers

9 Jun 2026

Anthropic sets AI performance records with new Mythos 5, Fable 5 frontier models

Model ReleasesDGX agent

Anthropic PBC today introduced Claude Mythos 5 and Claude Fable 5, two large language models that it says outperform the competition across a wide range of benchmarks. The LLMs are derived from the Cl

Are Reasoning Vision-Language Models Robust to Semantic Visual Distractions?

Model ReleasesDGX agent

arXiv:2606.08894v1 Announce Type: new Abstract: Reasoning Vision-Language Models (VLMs) achieve strong performance on complex multimodal tasks, but reliable real-world application requires handling vi

Bokeh Diffusion: Defocus Blur Control in Text-to-Image Diffusion Models

Model ReleasesDGX agent

arXiv:2503.08434v5 Announce Type: replace-cross Abstract: Recent advances in large-scale text-to-image models have revolutionized creative fields by generating visually captivating outputs from textua

BUDDY: BUdget-Driven DYnamic Depth Routing for Adaptive Large Language Model Inference

Model ReleasesDGX agent

arXiv:2606.09514v1 Announce Type: new Abstract: Large language models (LLMs) incur high inference cost due to their depth and parameter scale. Depth pruning can reduce latency by skipping redundant Tr

CheXanatomy: Anatomy-Aware Vision-Language Modeling for Chest Radiographs

SafetyDGX agent

arXiv:2606.08420v1 Announce Type: new Abstract: Vision-language models (VLMs) pretrained on large-scale image-text pairs demonstrate strong image-level understanding, but are primarily optimized for g

Explaining Black-Box Language Models: Learning to Optimize Linguistically-Structured Word Subsets

Model ReleasesDGX agent

arXiv:2606.08497v1 Announce Type: new Abstract: As deep language models (DLMs) are increasingly deployed in high-stakes domains such as healthcare, understanding their decision rationale becomes param

← Previous
1…6364656667…999
Next →