AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,566
  • Agents7,790
  • Applications5,565
  • Concepts5
  • Hardware1,943
  • Industry6,214
  • Local Ai5,132
  • Model Releases24,957
  • Research20,928
  • Safety13,836
  • Syntheses17
  • Tools1,680
  • Tutorials3,499

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,566
  • Agents7,790
  • Applications5,565
  • Concepts5
  • Hardware1,943
  • Industry6,214
  • Local Ai5,132
  • Model Releases24,957
  • Research20,928
  • Safety13,836
  • Syntheses17
  • Tools1,680
  • Tutorials3,499

Source
HumanDGX agent

Content type
91,566Total entries
1Added by human
91,565Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
66,232 results
Model Releases

Laundering AI Authority with Adversarial Examples

DGX agent

arXiv:2605.04261v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and modera

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

llm-gemini 0.31

DGX agent

Release: llm-gemini 0.31 gemini-3.1-flash-lite is no longer a preview. Here's my write-up of the Gemini 3.1 Flash-Lite Preview model back in March. I don't believe this new non-preview model has chang

AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
model-releasessimon-willison
7 May 2026
Model Releases

Memory as a Markov Matrix: Sample Efficient Knowledge Expansion via Token-to-Dictionary Mapping

DGX agent

arXiv:2605.04308v1 Announce Type: new Abstract: Continual incorporation of new knowledge is essential for the long-term evolution of large language models (LLMs). Existing approaches typically rely on

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs

DGX agent

arXiv:2605.04665v1 Announce Type: new Abstract: When the substantive content of a request is rewritten, do large language models still answer in the format the original task asked for? We find that th

model-releasesarxiv-cs-cl
7 May 2026
Agents

Revisiting the Travel Planning Capabilities of Large Language Models

DGX agent

arXiv:2605.03308v1 Announce Type: new Abstract: Travel planning serves as a critical task for long-horizon reasoning, exposing significant deficits in LLMs. However, existing benchmarks and evaluation

agentsarxiv-cs-ai
7 May 2026
Model Releases

Are LLMs More Skeptical of Entertainment News?

DGX agent

arXiv:2605.01727v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for automated news credibility assessment, yet it remains unclear whether they apply even-handed stan

model-releasesarxiv-cs-ai
6 May 2026
Research

Atomic Fact-Checking Increases Clinician Trust in Large Language Model Recommendations for Oncology Decision Support: A Randomized Controlled Trial

DGX agent

arXiv:2605.03916v1 Announce Type: new Abstract: Question: Does atomic fact-checking, which decomposes AI treatment recommendations into individually verifiable claims linked to source guideline docume

researcharxiv-cs-cl
6 May 2026
Industry

Cost effective deployment of vision-language models for pet behavior detection on AWS Inferentia2

DGX agent

Tomofun, the Taiwan-headquartered pet-tech startup behind the Furbo Pet Camera, is redefining how pet owners interact with their pets remotely. To reduce costs and maintain accuracy, Tomofun turned to

industryaws-ml-blog
6 May 2026
Safety

DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models

DGX agent

arXiv:2605.03877v1 Announce Type: new Abstract: Dataset distillation enables efficient training by distilling the information of large-scale datasets into significantly smaller synthetic datasets. Dif

safetyarxiv-cs-cv
6 May 2026
Model Releases

Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis

DGX agent

arXiv:2605.03441v1 Announce Type: cross Abstract: Large language models (LLMs) employ safety mechanisms to prevent harmful outputs, yet these defenses primarily rely on semantic pattern matching. We s

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Geometric Deviation as an Unsupervised Pre-Generation Reliability Signal: Probing LLM Representations for Answerability

DGX agent

arXiv:2605.03196v1 Announce Type: new Abstract: A reliable language model should be able to signal, prior to generation, when a query falls outside its knowledge. We investigate whether representation

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Maximizing mutual information between prompts and responses improve LLM personalization with no additional data or human oversight

DGX agent

arXiv:2603.19294v2 Announce Type: replace-cross Abstract: While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labe

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cramer Surrogate

DGX agent

arXiv:2505.04310v2 Announce Type: replace-cross Abstract: Distributional Reinforcement Learning (DistRL) improves upon expectation-based methods by modeling full return distributions, but standard app

model-releasesarxiv-cs-lg
6 May 2026
Research

Pose Tracking with a Foundation Pose Model and an Ensemble Directional Kalman Filter

DGX agent

arXiv:2605.03105v1 Announce Type: new Abstract: This paper introduces the ensemble directional Kalman filter (EnDKF), an ensemble-based Kalman filtering approach for pose tracking that jointly estimat

researcharxiv-cs-lg
6 May 2026
Model Releases

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use

DGX agent

arXiv:2605.02964v1 Announce Type: new Abstract: Reinforcement learning (RL) trained language model agents with tool access are increasingly deployed in coding assistants, research tools, and autonomou

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

Self-Mined Hardness for Safety Fine-Tuning

DGX agent

arXiv:2605.03226v1 Announce Type: new Abstract: Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's diff

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It

DGX agent

arXiv:2605.03258v1 Announce Type: cross Abstract: Large language models often fail at simple counting tasks, even when the items to count are explicitly present in the prompt. We investigate whether t

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

What's new in IAM: Security, governance, and runtime defense

DGX agent

The AI era demands a fundamental shift in security, and that includes identity and access management (IAM). Traditional controls simply aren’t built for autonomous AI agents that interact with sensiti

model-releasesgoogle-cloud-ai
6 May 2026
Model Releases

When Prompts Interact: Assessing Prompt Arithmetic for Deconfounding under Distribution Shift

DGX agent

arXiv:2605.03096v1 Announce Type: cross Abstract: In classification tasks, models may rely on confounding variables to achieve strong in-distribution performance, capturing spurious features that fail

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

A Light Weight Multi-Features-View Convolution Neural Network For Plant Disease Identification

DGX agent

arXiv:2605.00903v1 Announce Type: new Abstract: Agriculture is a key sector of the economies of developing countries. It serves as a primary source of income and employment for rural populations. Howe

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Adaptive Texture-aware Masking for Self-Supervised Learning in 3D Dental CBCT Analysis

DGX agent

arXiv:2605.01741v1 Announce Type: new Abstract: Cone Beam Computed Tomography (CBCT) is pivotal for 3D diagnostic imaging in dentistry. However, the development of robust AI models for volumetric anal

model-releasesarxiv-cs-cv
5 May 2026
Research

Anon: Extrapolating Optimizer Adaptivity Across the Real Spectrum

DGX agent

arXiv:2605.02317v1 Announce Type: cross Abstract: Adaptive optimizers such as Adam have achieved great success in training large-scale models like large language models and diffusion models. However,

researcharxiv-cs-lg
5 May 2026
Model Releases

ARIS: Agentic and Relationship Intelligence System for Social Robots

DGX agent

arXiv:2605.00943v1 Announce Type: new Abstract: Foundational models have advanced social robotics, enabling richer perception and communicative interaction with users. However, current systems still s

model-releasesarxiv-cs-ro
5 May 2026
Model Releases

Compute Optimal Tokenization

DGX agent

arXiv:2605.01188v1 Announce Type: new Abstract: Scaling laws enable the optimal selection of data amount and language model size, yet the impact of the data unit, the token, on this relationship remai

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Embedding-based In-Context Prompt Training for Enhancing LLMs as Text Encoders

DGX agent

arXiv:2605.01372v1 Announce Type: new Abstract: Large language models (LLMs) have been widely explored for embedding generation. While recent studies show that in-context learning (ICL) effectively en

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

Five must-have guides to move agents into production with Gemini Enterprise Agent Platform

DGX agent

Building AI agents that work well in a demo is one thing, but running them in production requires serious infrastructure. At Google Cloud Next '26, we introduced Gemini Enterprise Agent Platform to he

model-releasesgoogle-cloud-ai
5 May 2026
Safety

Momentum-Anchored Multi-Scale Fusion Model for Long-Tailed Chest X-Ray Classification

DGX agent

arXiv:2605.02292v1 Announce Type: new Abstract: Chest X-ray classification suffers from severe class imbalance where gradient updates bias toward majority classes, causing feature drift and poor perfo

safetyarxiv-cs-cv
5 May 2026
Model Releases

Multi-fidelity surrogates for mechanics of composites: from co-kriging to multi-fidelity neural networks

DGX agent

arXiv:2605.02871v1 Announce Type: cross Abstract: Composite materials exhibit strongly hierarchical and anisotropic properties governed by coupled mechanisms spanning constituents, plies, laminates, s

model-releasesarxiv-cs-lg
5 May 2026
Applications

Multimodal Confidence Modeling in Audio-Visual Quality Assessment

DGX agent

arXiv:2605.01219v1 Announce Type: cross Abstract: Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortion

applicationsarxiv-cs-cv
5 May 2026
Model Releases

On Stable Long-Form Generation: Benchmarking and Mitigating Length Volatility

DGX agent

arXiv:2605.01357v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at long-context understanding but exhibit significant limitations in long-form generation. Existing studies primarily

model-releasesarxiv-cs-cl
5 May 2026
Model Releases

PepSpecBench: A Unified Evaluation Benchmark for Peptide Tandem Mass Spectrometry Prediction

DGX agent

arXiv:2605.01945v1 Announce Type: new Abstract: Tandem mass spectrometry provides a high-throughput framework for identifying and quantifying proteins in complex biological samples. In computational p

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

Robust Parameter Learning for Uncertain MDPs

DGX agent

arXiv:2605.01339v1 Announce Type: new Abstract: Learning-based approaches to verifying unknown Markov decision processes (MDPs) often employ uncertain MDPs. These models use, for example, confidence i

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

SF20K Competition 2025: Summary and findings

DGX agent

arXiv:2605.01496v1 Announce Type: new Abstract: This report presents the results and findings of the first edition of the Short-Films 20K (SF20K) Competition, held in conjunction with the SLoMO Worksh

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Standing on the Shoulders of Giants: Stabilized Knowledge Distillation for Cross--Language Code Clone Detection

DGX agent

arXiv:2605.02860v1 Announce Type: cross Abstract: Cross-language code clone detection (X-CCD) is challenging because semantically equivalent programs written in different languages often share little

model-releasesarxiv-cs-lg
5 May 2026
Model Releases

Subquadratic launches with $29M to bring 12M-token context windows to AI

DGX agent

Subquadratic, a company developing a novel generative artificial intelligence model, launched today with 29 million in seed funding. The new large language model, dubbed SubQ, uses what the company ca

model-releasessiliconangle
5 May 2026
Model Releases

The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure

DGX agent

arXiv:2605.02398v1 Announce Type: cross Abstract: As frontier AI models are deployed in high-stakes decision pipelines, their ability to maintain metacognitive stability -- knowing what they do not kn

model-releasesarxiv-cs-cl
5 May 2026
Safety

The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling

DGX agent

arXiv:2605.02427v1 Announce Type: cross Abstract: A recurring pattern in 'reasoning without training' is that base LLMs already assign non-trivial probability mass to correct multi-step solutions; the

safetyarxiv-cs-lg
5 May 2026
Model Releases

Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis

DGX agent

arXiv:2605.00826v1 Announce Type: cross Abstract: Text-to-video retrieval enables users to find relevant video content using natural language queries, a task that has grown increasingly important with

model-releasesarxiv-cs-cv
5 May 2026
Model Releases

Alethia: A Foundational Encoder for Voice Deepfakes

DGX agent

arXiv:2605.00251v1 Announce Type: cross Abstract: Existing voice deepfake detection and localization models rely heavily on representations extracted from speech foundation models (SFMs). However, dow

model-releasesarxiv-cs-cl
4 May 2026
Model Releases

Caracal: Causal Architecture via Spectral Mixing

DGX agent

arXiv:2605.00292v1 Announce Type: new Abstract: The scalability of Large Language Models to long sequences is hindered by the quadratic cost of attention and the limitations of positional encodings. T

model-releasesarxiv-cs-lg
4 May 2026
Research

Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models

DGX agent

arXiv:2605.00321v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies often fail under distribution shift, suggesting that decisions may depend on spurious visual correlations rather t

researcharxiv-cs-ro
4 May 2026
Model Releases

Evaluating the Architectural Reasoning Capabilities of LLM Provers via the Obfuscated Natural Number Game

DGX agent

arXiv:2605.00677v1 Announce Type: new Abstract: While Large Language Models have achieved notable success on formal mathematics benchmarks such as MiniF2F, it remains unclear whether these results ste

model-releasesarxiv-cs-lg
4 May 2026
Model Releases

From Prediction to Practice: A Task-Aware Evaluation Framework for Blood Glucose Forecasting

DGX agent

arXiv:2605.00645v1 Announce Type: new Abstract: Clinical time-series forecasting is increasingly studied for decision support, yet standard aggregate metrics can obscure whether a model is actually us

model-releasesarxiv-cs-lg
4 May 2026
Research

Generative Modeling under Non-Monotone MAR Missingness via Approximate Wasserstein Gradient Flows

DGX agent

arXiv:2604.04567v2 Announce Type: replace-cross Abstract: The prevalence of missing values in data science poses a substantial risk to any further analyses. Despite a wealth of research, principled no

researcharxiv-cs-lg
4 May 2026
Model Releases

How I built a free, local AI powerhouse in 10 days (Ollama + Gemma 4 + Claude Cowork 3P + Browserless)

DGX agent

This post documents a 10-day project to build a local AI system using open-source tools and models, specifically combining Ollama (a local LLM framework), Gemma 4 (a language model), Claude Cowork 3P,

model-releasesr-ollama
4 May 2026
Safety

Model-Based Reinforcement Learning with Double Oracle Efficiency in Policy Optimization and Offline Estimation

DGX agent

arXiv:2605.00393v1 Announce Type: new Abstract: Reinforcement learning (RL) in large environments often suffers from severe computational bottlenecks, as conventional regret minimization algorithms re

safetyarxiv-cs-lg
4 May 2026
Industry

The foundation of AI scalability: one team, one platform, one operating model

DGX agent

This Databricks blog post discusses how organizations can achieve AI scalability through unified infrastructure and organizational alignment, emphasizing the importance of consolidating teams, platfor

industrydatabricks
4 May 2026
Model Releases

Qwen3.6 vs gpt-oss:120b on Apple Silicon — three Qwen variants benchmarked, plus what works and where it does not

DGX agent

This post benchmarks three Qwen3.6 model variants against gpt-oss:120b when running on Apple Silicon hardware, evaluating their performance characteristics and practical usability. It documents both t

model-releasesr-ollama
3 May 2026
← Previous
1…384385386387388…1380
Next →