AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,092 results
12 Aug 2026

From #22 to #4 on Legal Research Bench! Solid progress for Qwen3.8-Max. Thanks for highlighting~✨

Model ReleasesDGX agent

From #22 to #4 on Legal Research Bench! Solid progress for Qwen3.8-Max. Thanks for highlighting~✨ Qwen 3.8 Max nearly doubled its score on Legal Research Bench in under three months, climbing from #22

From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents

Model ReleasesDGX agent

arXiv:2608.10502v1 Announce Type: new Abstract: Persistent memory lets language-model agents reuse information across sessions, but it also makes errors durable: a poisoned, stale, or misattributed re

From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2608.10444v1 Announce Type: cross Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains. This pro

From Sync APIs to support for the GPT-5 model series and agentic workflows: What’s new in Azure Content Understanding – August 2026

Model ReleasesDGX agent

Enterprise content is no longer just something people read. AI apps and agents are only as useful as the information they can understand, yet much of the world’s enterprise knowledge is locked in docu

Gemma 4 QAT handles KV cache quantization MUCH better, KLD benchmarks show

Model ReleasesDGX agent

Link to the article: KV Cache Quantization on Gemma 4 31B: Non-QAT vs QAT KLD benchmarks with BeeLlama.cpp v0.4.3, fork of llama.cpp with more KV cache quantization options, comparing Gemma Q4_0 non-Q

GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning

Model ReleasesDGX agent

arXiv:2608.10494v1 Announce Type: new Abstract: Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challen

GeoSeg-OV: Bridging Geospatial Gaps with Structural Guidance for Open-Vocabulary Remote Sensing Segmentation

Model ReleasesDGX agent

arXiv:2608.10426v1 Announce Type: new Abstract: Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories sp

GESTO: Human-Centric Spatio-Temporal Memory for Reasoning in Dynamic Scenes

Model ReleasesDGX agent

arXiv:2608.10886v1 Announce Type: new Abstract: Robots operating in human environments need memories that capture not only what objects exist and where, but also how people use them over time and how

give 4.6 a try and let us know how it goes. your feedback is a big part of why the model gets better with each iteration.

Model ReleasesDGX agent

give 4.6 a try and let us know how it goes. your feedback is a big part of why the model gets better with each iteration. SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, j

GLAM: Efficient Continual Learning at Scale via Grouped LoRA Adapter Merging

Model ReleasesDGX agent

arXiv:2509.13211v4 Announce Type: replace Abstract: The ability to learn continuously over time remains a major challenge for modern machine learning systems, even in the era of Foundation Models. Whi

Google DeepMind launches SL2T, a multilingual sign-language-to-text model debuting on the Pixel 11 in Gboard and Live Transcribe, first with ASL and English (Mike Wheatley/SiliconANGLE)

Model ReleasesDGX agent

Mike Wheatley / SiliconANGLE: Google DeepMind launches SL2T, a multilingual sign-language-to-text model debuting on the Pixel 11 in Gboard and Live Transcribe, first with ASL and English — Google Deep

Google unveils the $399 Pixel Watch 5 with a satin pyrite case finish, offline Gemini, proactive AI suggestions, better GPS maps, and insulin resistance trends (Victoria Song/The Verge)

Model ReleasesDGX agent

Victoria Song / The Verge: Google unveils the 399 Pixel Watch 5 with a satin pyrite case finish, offline Gemini, proactive AI suggestions, better GPS maps, and insulin resistance trends — The 399 Goog

Google unveils the 899+ Pixel 11, 1,099+ 11 Pro, and $1,299+ 11 Pro XL, with a Tensor G6, new Gemini features, Magic Capture to pick the best frames, and more (Ivan Mehta/TechCrunch)

Model ReleasesDGX agent

Ivan Mehta / TechCrunch: Google unveils the 899+ Pixel 11, 1,099+ 11 Pro, and $1,299+ 11 Pro XL, with a Tensor G6, new Gemini features, Magic Capture to pick the best frames, and more — For the last f

Grok 4.6 is an excellent model. I’ve been using it heavily for the past couple of weeks and it handles everything from simple coding & code …

Model ReleasesDGX agent

Grok 4.6 is an excellent model. I’ve been using it heavily for the past couple of weeks and it handles everything from simple coding & code review all the way to designing and debugging complex system

Grok 4.6 is now one of the top models in the world for agentic workflows It ranks #1 on the Artificial Analysis Agentic Index, tied with Cla…

Model ReleasesDGX agent

Grok 4.6 is now one of the top models in the world for agentic workflows It ranks #1 on the Artificial Analysis Agentic Index, tied with Claude Opus 5 Max Agentic AI is about more than answering quest

Grok 4.6 is objectively #1 when considering intelligence, speed & cost

Model ReleasesDGX agent

Grok 4.6 is objectively #1 when considering intelligence, speed & cost SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with

Grok 4.6 reaches 1753 ELO

Model ReleasesDGX agent

Grok 4.6 topped the GDPVal-AA benchmark with an Elo score of 1,753. It surpassed competitors Fable 5 Max (1,741 Elo), GPT‑5.6 Sol Max (1,728 Elo) and Grok 4.5 High (1,526 Elo). Elon Musk publicly ackn

Grok Bot

Model ReleasesDGX agent

Grok Bot Here's my Grok Bot team: - Webby: Web designer - Shotry: Short-form content creator - Writey: Article/Newsletter writer - Claude Code: Grok agent that specializes in CC - Codex: Same as the a

Hidden Reasoning from Claude and GPT are Decoded, and it is interesting

Model ReleasesDGX agent

Yesteday a paper showed a gap that allows to see 100% of the reasoning tokens form ALL Claude and GPT models Stealing Reasoning Traces from Proprietary LLM APIs. check it out, they have published lots

HNDiff: Haze-Noise Diffusion for Image Dehazing

Model ReleasesDGX agent

arXiv:2608.10995v1 Announce Type: new Abstract: Existing diffusion-based methods have recently made significant progress in image dehazing. However, they typically neglect the physics of haze formatio

HoosierHelp: Benchmarking LLM Agents for Social Service Navigation

Model ReleasesDGX agent

arXiv:2608.09946v1 Announce Type: cross Abstract: Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM ag

How OneAdvanced deployed over 50 AI agents on UK-sovereign AWS

Model ReleasesDGX agent

Learn how OneAdvanced, a UK enterprise software provider, built a UK-sovereign AI platform by self-hosting Llama 4 Maverick and Llama Guard 4 on Amazon SageMaker AI, with a RAG pipeline on pgvector an

How Robust Are LLMs to Vietnamese Dialects?

Model ReleasesDGX agent

arXiv:2608.10414v1 Announce Type: new Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects th

HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2506.03922v4 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benchma

HUI360: A 360{eg} Egocentric Dataset and Baselines for Human-Robot Interaction Anticipation

Model ReleasesDGX agent

arXiv:2608.11051v1 Announce Type: new Abstract: As robots increasingly operate in human-populated environments, anticipating human intentions is essential for enabling proactive and socially aware beh

HyperShape: Hyperelasticity Across Diverse Shapes

Model ReleasesDGX agent

arXiv:2608.09938v1 Announce Type: cross Abstract: Hyperelastic deformations are highly sensitive to domain geometry and boundary conditions, making generalization across both a critical capability for

Idea for a deepseek-v4-flash-0731 backed automated research workflow to be leveraged via qwen3.6/3.8 27b for difficult tasks that require highly technical, not easy to find information.

Model ReleasesDGX agent

Sometimes you have tasks that are outside of your expertise and the idea is this workflow automation could be leveraged to manage to have local AI figure it out using research from his workflow gather

imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here …

Model ReleasesDGX agent

imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here to top it at a great price, congrats to @SpaceXAI team & loo

ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover

Model ReleasesDGX agent

arXiv:2608.10545v1 Announce Type: cross Abstract: Edge LLMs must preserve inference continuity when a user hands over between edge nodes, requiring key-value (KV) cache transfer to the target node. Ho

Implicit representations are dead. Long live explicit primitives!

Model ReleasesDGX agent

arXiv:2608.10001v1 Announce Type: cross Abstract: Continuous parameterization of medical data has emerged as a powerful paradigm for resolution-independent image representation. While Implicit Neural

InSight-doc: Agentic Visual Perception for Long-Document Understanding

Model ReleasesDGX agent

arXiv:2608.10628v1 Announce Type: cross Abstract: Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we

Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show ope…

Model ReleasesDGX agent

Interesting research suggests caution in determining which AI company is winning by looking at any one source.. OpenRouter seems to show open weights winning over time, but work submitted to Pangram i

Interpreting Language Model Hidden States at Scale

Model ReleasesDGX agent

arXiv:2608.10260v1 Announce Type: new Abstract: Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions d

Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability

Model ReleasesDGX agent

arXiv:2608.10172v1 Announce Type: new Abstract: Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or

Introspective Attention Modulation for Safe Text-to-Image Generation

Model ReleasesDGX agent

arXiv:2607.14945v2 Announce Type: replace Abstract: State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Pr

Invertible Logits Transformation for Accuracy-Preserving Post-Hoc Uncertainty Calibration

Model ReleasesDGX agent

arXiv:2608.10372v1 Announce Type: new Abstract: Post-hoc calibration aligns a classifier's predicted confidences with its empirical accuracy without retraining. An ideal calibrator should correct nonl

Is There Really a Camouflaged Object? Towards Realistic Camouflaged Object Detection

Model ReleasesDGX agent

arXiv:2608.11135v1 Announce Type: new Abstract: Camouflaged object detection (COD) aims to segment objects that are visually concealed in their surroundings and has attracted increasing attention in r

Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility

Model ReleasesDGX agent

arXiv:2608.10315v1 Announce Type: cross Abstract: Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or s

KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning

Model ReleasesDGX agent

arXiv:2501.11655v3 Announce Type: replace-cross Abstract: This paper proposes a novel learning approach for designing Kazantzis-Kravaris or nonlinear Luenberger (KKL) observers for autonomous nonlinea

Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies

Model ReleasesDGX agent

arXiv:2608.10273v1 Announce Type: cross Abstract: Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of transmitting p

Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability

Model ReleasesDGX agent

arXiv:2608.10300v1 Announce Type: new Abstract: Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human re

Long-Time Trajectory Approximation via SA-NODEs: Model Predictive and Floquet Strategies

Model ReleasesDGX agent

arXiv:2608.10738v1 Announce Type: new Abstract: We study the approximation of dynamical systems by semi-autonomous neural ordinary differential equations (SA-NODEs) over long time horizons. For a sing

MAD-HOI: Masked Autoregressive Diffusion for Generating Articulated Hand Object Interactions from Text

Model ReleasesDGX agent

arXiv:2608.10162v1 Announce Type: new Abstract: Methods for text-based generation of hand-object interaction (HOI) sequences primarily focus on producing smooth, physically plausible trajectories. A t

MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows

Model ReleasesDGX agent

arXiv:2608.10509v1 Announce Type: new Abstract: Shared memory helps language-model agents reuse information across long workflows, yet relevant evidence may not be admissible for a particular agent or

Mapping and Measuring the Behavioral Evolution of Large Language Models

Model ReleasesDGX agent

arXiv:2608.11027v1 Announce Type: cross Abstract: Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across genera

Masked Neural Detection for Run-Length-Limited Channel Coding in Molecular Communication

Model ReleasesDGX agent

arXiv:2606.12489v2 Announce Type: replace-cross Abstract: Molecular communication (MC) suffers from severe diffusion memory because molecules released for one symbol may arrive during later symbol int

Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems

Model ReleasesDGX agent

arXiv:2607.10309v2 Announce Type: replace Abstract: Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Things (AIoT

Measuring Semantic Abstractness of SAE Features via Nonlocality

Model ReleasesDGX agent

arXiv:2608.10537v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the c

MindTopo reveals VLMs’ spatial reasoning abilities

Model ReleasesDGX agent

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post

Mistral says its platform will support third-party open models, starting with Z.ai's GLM-5.2, and run them on the same infrastructure as its own models (Mistral AI Blog)

Model ReleasesDGX agent

Mistral AI Blog: Mistral says its platform will support third-party open models, starting with Z.ai's GLM-5.2, and run them on the same infrastructure as its own models — At Mistral, we believe every

More Accurate, Less Human: Gestalt Grouping in Vision Models

Model ReleasesDGX agent

arXiv:2608.10195v1 Announce Type: new Abstract: Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into r

Multimodal Item Parameter Estimation using Simulated Response Probabilitie

Model ReleasesDGX agent

arXiv:2608.10154v1 Announce Type: cross Abstract: We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large

myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR

Model ReleasesDGX agent

arXiv:2608.11036v1 Announce Type: new Abstract: Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work prese

Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents

Model ReleasesDGX agent

arXiv:2608.09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure be

Neural Introspection Gating for Adaptive KV-Cache Reuse in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2608.10824v1 Announce Type: cross Abstract: Vision-Language-Action(VLA) models map camera images and language instructions directly to motor commands through a single autoregressive transformer.

New Muse-Glimmer-30B SoTA Quants - hopefully a new lineup :)

Model ReleasesDGX agent

Hey Folks, I've been making quants for a while - recently I took a short break to get into hardcore research (submitted my first EMNLP paper during it!). Along the way, I built up a little arsenal of

No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Model ReleasesDGX agent

arXiv:2503.05061v3 Announce Type: replace Abstract: Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as bus

NullEdit: Stealthy Image Protection via VLM Condition Redirection

Model ReleasesDGX agent

arXiv:2608.10870v1 Announce Type: new Abstract: Modern image editors combine vision-language models (VLMs) with diffusion transformer backbones to modify a single reference image according to instruct

Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research

Model ReleasesDGX agent

arXiv:2608.10363v1 Announce Type: new Abstract: AI agents can accelerate nutrition research, but their analyses inherit the identity, semantic, and release ambiguities of the underlying data. We prese

NVIDIA CEO Tops Glassdoor’s 2026 List of Best CEOs

Model ReleasesDGX agent

NVIDIA founder and CEO Jensen Huang is ranked No. 1 on Glassdoor’s Best CEOs list for 2026. In the just-released ranking, recognition is earned directly from the people who know their leadership the b

← Previous
12345…369
Next →