AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,428 results
7 May 2026

Gemini 3.1 Flash-Lite is now generally available on Gemini Enterprise

Model ReleasesDGX agent

Today, we’re thrilled to announce that Gemini 3.1 Flash-Lite, our fastest and most cost-efficient Gemini 3 series model yet, is now generally available. Designed for ultra-low latency, high-volume tas

5 May 2026

Agent Factory Recap: How Gemma 4 Taught Itself Physics

Model ReleasesDGX agent

In this episode of The Agent Factory, Vlad Kolesnikov and I sat down with Omar Sanseviero from the Developer Experience team at Google DeepMind. We explored the groundbreaking release of Gemma 4: a ne

22 Apr 2026

What’s next in Google AI infrastructure: Scaling for the agentic era

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

AI is evolving from answering questions to reasoning and taking action. Companies who want to lead in today’s agentic era require computing infrastructure designed and optimized for these new requirem

21 Apr 2026

A Probabilistic Consensus-Driven Approach for Robust Counterfactual Explanations

Model ReleasesDGX agent

arXiv:2604.17494v1 Announce Type: new Abstract: Counterfactual explanations (CFEs) are essential for interpreting black-box models, yet they often become invalid when models are slightly changed. Exis

ESsEN: Training Compact Discriminative Vision-Language Transformers in a Low-Resource Setting

Model ReleasesDGX agent

arXiv:2604.18452v1 Announce Type: cross Abstract: Vision-language modeling is rapidly increasing in popularity with an ever expanding list of available models. In most cases, these vision-language mod

User-Assistant Bias in LLMs

Model ReleasesDGX agent

arXiv:2508.15815v3 Announce Type: replace Abstract: Modern large language models (LLMs) are typically trained and deployed using structured role tags (e.g. system, user, assistant, tool) that explicit

14 Apr 2026

SLM Finetuning for Natural Language to Domain Specific Code Generation in Production

Model ReleasesDGX agent

arXiv:2604.09952v1 Announce Type: new Abstract: Many applications today use large language models for code generation; however, production systems have strict latency requirements that can be difficul

What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction

Model ReleasesDGX agent

arXiv:2407.08101v4 Announce Type: replace Abstract: Vision-language models have shown impressive progress in recent years. However, existing models are largely limited to turn-based interactions, wher

13 Apr 2026

EXAONE 4.5 Technical Report

Model ReleasesDGX agent

arXiv:2604.08644v1 Announce Type: new Abstract: This technical report introduces EXAONE 4.5, the first open-weight vision language model released by LG AI Research. EXAONE 4.5 is architected by integr

10 Apr 2026

Gemma 4:e4b offloads to RAM despite having just half of VRAM used.

Model ReleasesDGX agent

Users on the r/ollama subreddit reported that the **Gemma 4 E4B** model in Ollama offloads layers to RAM even when GPU VRAM is only partially utilized. This behavior is linked to how Ollama and lla...

14 Aug 2026

A Bayes-Markov Neuromorphic Model of Cortical Orientation Selectivity: A Computational Re-implementation and Quantitative Simulation Study

Model ReleasesDGX agent

arXiv:2608.12388v1 Announce Type: cross Abstract: The emergence of orientation selectivity in the primary visual cortex (V1) remains a central question in computational neuroscience. Shirazi's Bayes-M

Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents

Model ReleasesDGX agent

arXiv:2608.12342v1 Announce Type: new Abstract: Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies ha

Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities

ResearchDGX agent

arXiv:2608.12374v1 Announce Type: cross Abstract: Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays a crucial r

Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds

Model ReleasesDGX agent

arXiv:2608.13069v1 Announce Type: new Abstract: Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically

FIRE-VLA: Failure-Informed Self-Evolution for Vision-Language-Action Models in Autonomous Driving

Model ReleasesDGX agent

arXiv:2608.13395v1 Announce Type: new Abstract: Reinforcement learning improves autonomous-driving vision-language-action (VLA) models by evaluating trajectories sampled from the current policy. Group

Jointly Predicting Courses and Grades Using a Transformer-Based Model

TutorialsDGX agent

arXiv:2608.13409v1 Announce Type: new Abstract: Existing predictive models in learning analytics often treat student academic history as a simple sequence, overlooking the concurrent nature of courses

Large Language Models Pass the History Exam But Miss the <<History>>: A Polish High School Exit Exam Matura Benchmark

Model ReleasesDGX agent

arXiv:2608.12343v1 Announce Type: new Abstract: AI chatbots are widely used by students as knowledge sources, yet LLM benchmarks rarely assess interpretative historical reasoning. We evaluate eight le

LFM 2.5 2.6B is the best small model for tool use I have ever used.

Local AiDGX agent

I'm working on a local perplexity/AI search comprised of a custom harness and a further trained model. LFM 2.6 is almost to spec with no additional training, it is very good. submitted by /u/thebadsli

Novels generated by language models show compressed formal variation

Model ReleasesDGX agent

arXiv:2608.12630v1 Announce Type: cross Abstract: While large language models can generate entire novels, there is little information about the level of formal variation in their output over many gene

Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts

ResearchDGX agent

arXiv:2608.12446v1 Announce Type: cross Abstract: Sleep stage classification is important for the diagnosis and management of sleep disorders, yet most automatic staging studies evaluate models agains

Scaling Automatic Research Agents via World Models

SafetyDGX agent

arXiv:2608.12564v1 Announce Type: new Abstract: Automating empirical research is a long-standing direction of AI. Recent automatic research (AutoResearch) agents bring this goal within reach, as moder

13 Aug 2026

Air Quality Station Simulation via LSTM and Attention-Based Modelling

Model ReleasesDGX agent

arXiv:2608.11839v1 Announce Type: new Abstract: Poor air quality in urban areas is driven by a complex chain of processes and presents a significant public health concern. To better understand and con

Can Vision Models Read the Radar Display? On the Feasibility of Radar Imagery for Air Traffic Complexity Estimation

ResearchDGX agent

arXiv:2608.11810v1 Announce Type: new Abstract: Air traffic controllers perceive traffic complexity through the radar display, suggesting that a computer vision model operating on the same imagery may

CT-DeltaBench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.11534v1 Announce Type: new Abstract: In medical imaging, the clinical value of Computed Tomography (CT) lies not only in depicting current disease status, but crucially in enabling longitud

How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?

ResearchDGX agent

arXiv:2603.00056v2 Announce Type: replace-cross Abstract: STEM Mental models can play a critical role in assessing students' conceptual understanding of a topic. They not only offer insights into what

Orientation, not magnitude: the causal structure of task-vector interference in merged language models

Model ReleasesDGX agent

arXiv:2608.11797v1 Announce Type: new Abstract: Model merging by task arithmetic works until it doesn't, and the field diagnoses why with magnitudes: layerwise representation bias, deviations from cro

Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop

Model ReleasesDGX agent

arXiv:2608.11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase

Qwen-MusicAVQA-7B: A Multimodal Model for Music Audio-Visual QA

Model ReleasesDGX agent

arXiv:2608.11329v1 Announce Type: cross Abstract: A common approach to adding audio to a vision-language model is to train or adapt a large omni-modal system. We show that a lightweight alternative ca

Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models

Model ReleasesDGX agent

arXiv:2608.11657v1 Announce Type: cross Abstract: We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization problem into

Writer launches major agentic AI improvements with Palmyra X6 flagship model

Model ReleasesDGX agent

Generative artificial intelligence startup Writer Inc. today announced the release of its next-generation flagship model, Palmyra X6, designed to deliver frontier-level performance for marketing and r

12 Aug 2026

Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift

Local AiDGX agent

arXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a eta-fraction of shifted tar

Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models

SafetyDGX agent

arXiv:2608.10393v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial r

HQ-DM: Single Hadamard Transformation-Based Quantization-Aware Training for Low-Bit Diffusion Models

ResearchDGX agent

arXiv:2512.05746v3 Announce Type: replace Abstract: Diffusion models have demonstrated significant applications in the field of image generation. However, their high computational and memory costs pos

Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies

Model ReleasesDGX agent

arXiv:2608.10273v1 Announce Type: cross Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting p

Palo Alto Networks to run OpenAI cyber models inside customer networks

IndustryDGX agent

Palo Alto Networks Inc. said today its Unit 42 consulting arm will put OpenAI Group PBC’s frontier cyber models to work inside customer environments, expanding a service it launched earlier this year

ReCBM: Uncertainty-Gated Relational Reasoning for Concept Bottleneck Models

TutorialsDGX agent

arXiv:2608.10004v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) provide an interpretable framework by grounding predictions in human-understandable concepts, enabling semantic inspect

Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning

SafetyDGX agent

arXiv:2608.11204v1 Announce Type: cross Abstract: Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g.,

11 Aug 2026

Activation Probes Surface Code-Security Signals that the Model's Output Misses

AgentsDGX agent

arXiv:2608.09643v1 Announce Type: cross Abstract: AI coding agents now write a growing share of production code, and human security review does not scale at the rate code is generated. The agents in w

An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography

Model ReleasesDGX agent

arXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency.

CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting

Model ReleasesDGX agent

arXiv:2608.07693v1 Announce Type: cross Abstract: Generative traffic video forecasting aims to synthesize long-horizon, temporally coherent future videos of traffic scenes from a short observation his

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects

ResearchDGX agent

arXiv:2608.08067v1 Announce Type: cross Abstract: Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scenarios due to

Ensemble learning of pathology foundation models for precision oncology

TutorialsDGX agent

arXiv:2508.16085v2 Announce Type: replace Abstract: Histopathology is essential for cancer diagnosis and treatment selection, and pathology foundation models learn visual representations from whole-sl

FemWear: A Specialized Wearable Foundation Model for Women's Health

Model ReleasesDGX agent

arXiv:2608.08244v1 Announce Type: new Abstract: General wearable foundation models are pretrained across broad sensor streams and populations, but are not designed around women's-health tasks. We intr

FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2608.08736v1 Announce Type: new Abstract: Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Models (MLLMs) i

From Independent to Correlated Diffusion: Generalized Generative Modeling with Probabilistic Computers

ResearchDGX agent

arXiv:2603.27996v2 Announce Type: replace Abstract: Diffusion models have emerged as a powerful framework for generative tasks in deep learning. They decompose generative modeling into two computation

Fusion Training for Mathematical Generalization in Large Language Models

Model ReleasesDGX agent

arXiv:2608.09893v1 Announce Type: cross Abstract: Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinking mode and

GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning

Local AiDGX agent

arXiv:2608.07619v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but often degrade under visual and environmental shifts. Latent worl

Local Benchmark : Muse Glimmer 30B vs Qwen 3.6 27B vs Gemma4 31B (and many other models and finetunes)

Model ReleasesDGX agent

Needs a lot of requests compared to Qwen (almost twice) and Gemma (almost x3). Final score is fine, even though it is 'not a coding model' https://wonderrico.github.io/local_llm_benchmark/benchmark-ma

☁️Mistral is bringing together the inference infrastructure, open models, and long-term commitments Europe needs to control its AI future, a…

Model ReleasesDGX agent

☁️Mistral is bringing together the inference infrastructure, open models, and long-term commitments Europe needs to control its AI future, and setting a roadmap for the world. 🧵: https://mistral.ai/ne

MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs

ResearchDGX agent

arXiv:2510.19366v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) scales model capacity through sparse activation, and is becoming an important architecture for large language models (LLMs)

MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

Model ReleasesDGX agent

arXiv:2603.28590v3 Announce Type: replace Abstract: Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When such a mis

Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models

Model ReleasesDGX agent

arXiv:2608.09227v1 Announce Type: new Abstract: Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference computationally pr

PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.09772v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason

RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

Model ReleasesDGX agent

arXiv:2608.08081v1 Announce Type: cross Abstract: Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneo

Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior

ResearchDGX agent

arXiv:2606.22790v2 Announce Type: replace-cross Abstract: In this paper, we investigate the tradeoffs between compute allocation and model performance for two speech processing tasks: Automatic Speech

Task-to-Model Optimization for Enterprise LLM Coding Assistants: A Data-Driven Framework for Cost-Optimal Routing

Model ReleasesDGX agent

arXiv:2608.08528v1 Announce Type: new Abstract: Enterprise AI coding assistants incur substantial inference spend, and naive token-cost minimization often fails to reduce end-to-end cost once retries,

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

Model ReleasesDGX agent

arXiv:2608.08650v1 Announce Type: new Abstract: Mixture-of-Experts models increase parameter capacity while keeping the computation activated by each token bounded, but their architectural evolution c

TokenPrint: A Calibrated Token-Space Fingerprint for Language-Model Provenance

ResearchDGX agent

arXiv:2608.08139v1 Announce Type: new Abstract: Establishing the provenance of a language model---including its base checkpoint and possible overlap in training distributions---is a governance challen

Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation

Model ReleasesDGX agent

arXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha

World Simulator: Queer Erotica and the Absurdity of AI Video Models That Promise the World

ResearchDGX agent

arXiv:2608.07510v1 Announce Type: cross Abstract: Increasingly, AI video models are marketed as 'world simulators,' suggesting their ability to model infinite realities. Despite such claims, these mod

← Previous
1…7677787980…1008
Next →