AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,084 results
12 Aug 2026

The model punches above its weight, outperforming Gemma 4 E2B and Ministral 3 3B across a broad range of visual understanding benchmarks and…

Model ReleasesDGX agent

The model punches above its weight, outperforming Gemma 4 E2B and Ministral 3 3B across a broad range of visual understanding benchmarks and delivers particularly strong results on document understand

The most dangerous document extraction failure isn't a wrong value. It's a missing row that looks like nothing is wrong. We released Extract…

Model ReleasesDGX agent

The most dangerous document extraction failure isn't a wrong value. It's a missing row that looks like nothing is wrong. We released ExtractBench yesterday: 370 enterprise docs, 14 systems. The hardes

The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degrada

The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages

Model ReleasesDGX agent

arXiv:2606.15821v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, fo

Today, we’re adding another member to our model family. Meet North Micro Vision. Our smallest vision-language model yet, ideal for sophistic…

Model ReleasesDGX agent

Today, we’re adding another member to our model family. Meet North Micro Vision. Our smallest vision-language model yet, ideal for sophisticated document understanding. Available open-source under an

Toward Human Rights Benchmarking for LLMs: A Pilot Methodology

Model ReleasesDGX agent

arXiv:2608.10268v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exis

Towards Efficient Reasoning in LLM-Based Recommender Systems via Model Merging

Model ReleasesDGX agent

arXiv:2608.10447v1 Announce Type: cross Abstract: Large language model-based recommender systems are increasingly adopting slow-thinking models that generate step-by-step reasoning before making predi

Towards Unified Dynamic Face Landmark Detection

Model ReleasesDGX agent

arXiv:2608.10346v1 Announce Type: cross Abstract: Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional limitations

TRACE: Trustworthy Retrieval-Augmented Conversational Engine

Model ReleasesDGX agent

arXiv:2608.10176v1 Announce Type: new Abstract: Public service chatbots are expected to deliver recommendations from an underlying public service directory, while also making sure that the recommendat

Try Grok 4.6 on tough real-world tasks!

Model ReleasesDGX agent

Try Grok 4.6 on tough real-world tasks! imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here to top it at a great

Try Qwen-Image-3.0 on @openart_ai! 🎨👀

Model ReleasesDGX agent

Try Qwen-Image-3.0 on @openart_ai! 🎨👀 Qwen Image 3.0 is now on OpenArt ✨ The most Real Qwen image model yet. Native text across 12 languages, precise 10px type, and full interfaces like web pages, gam

Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification

Model ReleasesDGX agent

arXiv:2608.10007v1 Announce Type: cross Abstract: The current state-of-the-art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep RVFL (edRVFL

UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs

Model ReleasesDGX agent

arXiv:2608.10042v1 Announce Type: cross Abstract: Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool u

UT-ACA: Uncertainty-Triggered Adaptive Context Allocation for Long-Context Inference

Model ReleasesDGX agent

arXiv:2603.18446v2 Announce Type: replace Abstract: Long-context inference remains challenging for large language models due to attention dilution and out-of-distribution degradation. Context selectio

V-FiLLM: Verified Financial LLM Reasoning Benchmark

Model ReleasesDGX agent

arXiv:2608.11047v1 Announce Type: new Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains compar

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?

Model ReleasesDGX agent

arXiv:2608.10875v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained re

VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

Model ReleasesDGX agent

arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world

Vision-Language-Motion Maps: An Open-Vocabulary, Uncertainty-Aware, Queryable Motion Attribute for 3D Scene Maps

Model ReleasesDGX agent

arXiv:2607.16173v2 Announce Type: replace Abstract: Open-vocabulary 3D maps let robots answer language queries about what and where, but they assume a static world and cannot answer queries about how

Visual Geometry Foundation-Aware Gaussians for Single-Frame Surround-View Driving Reconstruction

Model ReleasesDGX agent

arXiv:2608.10682v1 Announce Type: new Abstract: Single-frame surround-view reconstruction faces severe geometric instability and rendering artifacts due to minimal inter-camera overlap. While existing

VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation

Model ReleasesDGX agent

arXiv:2608.10359v1 Announce Type: cross Abstract: As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, l

We are in an insane run of open-weight drops. Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: …

Model ReleasesDGX agent

We are in an insane run of open-weight drops. Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B

We have her dash camera which shows they are lying. The agents are wearing body cameras and should have dash cameras of their own. If what t…

Model ReleasesDGX agent

We have her dash camera which shows they are lying. The agents are wearing body cameras and should have dash cameras of their own. If what they say happened was true they wouldn’t be issuing statement

We wrote a 36-page ArXiv whitepaper on ExtractBench 🧑‍🔬 , our effort to create the most comprehensive, schema-guided, real-world document …

Model ReleasesDGX agent

We wrote a 36-page ArXiv whitepaper on ExtractBench 🧑‍🔬 , our effort to create the most comprehensive, schema-guided, real-world document extraction benchmark. It’s extremely detailed and covers every

What unique, custom QOL upgrades have you given your local agents?

Model ReleasesDGX agent

Warning: Kinda long post. If you don't like reading, please skip for your own sanity. Also, I've got nothing to sell, just a tinkerer, so I just want to share ideas and learn from you guys too. When I

When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

Model ReleasesDGX agent

arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of

When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.11024v1 Announce Type: new Abstract: Attribute hallucination---where vision-language models (VLMs) correctly identify an object but mischaracterize its properties---is prevalent yet mechani

Whole-Body Planning for Humanoids Navigating Confined Spaces via Self-Collision Avoidance References

Model ReleasesDGX agent

arXiv:2608.10220v1 Announce Type: new Abstract: Humanoid locomotion in highly confined environments requires navigating dense environmental obstacles and complex self-collision bounds while maintainin

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

Model ReleasesDGX agent

arXiv:2608.11095v1 Announce Type: new Abstract: Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wh

Withholding the Completing Chunk: Deterministic Pair-Completion Guardrails for Streaming LLM Output

Model ReleasesDGX agent

arXiv:2608.10279v1 Announce Type: cross Abstract: Streaming language-model output creates a release-timing problem: complete-response moderation acts after streamed text has escaped, whereas repeated

Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data

Model ReleasesDGX agent

arXiv:2608.11022v1 Announce Type: cross Abstract: Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their con

11 Aug 2026

1 Day in and I feel okay saying Muse-Glimmer-30B finally beats 3.6-27B for the size in some use-cases

Model ReleasesDGX agent

A few things right off the bat: it reasons very efficiently. Like Grok 4.5 levels of efficient thinking it quantizes very well. My first few tests with iq3_xxs were better than Qwen/Gemma behaved at t

10 year garbage card for local llms

Model ReleasesDGX agent

Hello everyone! ​I like dumb things. I like working with weak computers and microcontrollers. I like the simplicity and low electricity usage. Simply put, the efficiency of a 'dumb' PC. ​The first tim

12GB VRAM gang, what's our plan?

Model ReleasesDGX agent

Seems like we're limited to qwen finetuned MoEs for now. Looking at the current landscape - focus seems to be on dense models (muse glimmer 30b, qwen 3.8 27b) for smaller setups. Is upgrading to 24GB

360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents

Model ReleasesDGX agent

arXiv:2608.08814v1 Announce Type: cross Abstract: We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment construc

⚡️A coalition that secures long-term AI capacity: We’re aggregating long-term compute demand in Europe to determine what capacity is built, …

Model ReleasesDGX agent

⚡️A coalition that secures long-term AI capacity: We’re aggregating long-term compute demand in Europe to determine what capacity is built, where it’s located, and whom it serves. Through these multi-

A Control Function Framework for Mitigating Position Bias in Learning to Rank Systems

Model ReleasesDGX agent

arXiv:2506.06989v3 Announce Type: replace-cross Abstract: Learning-to-rank (LTR) systems commonly depend on implicit feedback, such as user clicks, because it is easy to collect and can serve as a val

A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences

Model ReleasesDGX agent

arXiv:2608.08240v1 Announce Type: new Abstract: This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage

A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence

Model ReleasesDGX agent

arXiv:2501.17629v2 Announce Type: replace-cross Abstract: Several studies claim that large language models have passed the Turing Test and hence can 'think', yet none follow Turing's original instruct

A Tight Lower Bound for Smooth Nonconvex Stochastic Optimization with Bounded Gradient Noise

Model ReleasesDGX agent

arXiv:2608.09004v1 Announce Type: cross Abstract: We prove a sharp lower bound for smooth nonconvex stochastic optimization with uniformly bounded gradient noise. In the (K=1) fresh-sample model, ever

A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents

Model ReleasesDGX agent

arXiv:2608.09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.

Accelerate PostgreSQL migrations using Gemini in Database Migration Service

Model ReleasesDGX agent

Imagine this scenario: Your team decides to migrate a core application from an existing commercial database like Oracle or SQL Server to open source PostgreSQL or a fully managed service such as Alloy

ACEvo: Adversarial Co-Evolution of Problem Distributions and Solvers for Combinatorial Optimization

Model ReleasesDGX agent

arXiv:2506.02594v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to synthesize heuristic programs, yet most existing pipelines optimize solvers against fixed benc

ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

Model ReleasesDGX agent

arXiv:2608.09476v1 Announce Type: cross Abstract: Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavio

ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception

Model ReleasesDGX agent

arXiv:2607.10180v2 Announce Type: replace-cross Abstract: We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception. T

AdaDINO: Pair-Aware In-Backbone Adaptation of Frozen DINO for Efficient Remote Sensing Change Detection

Model ReleasesDGX agent

arXiv:2608.07982v1 Announce Type: new Abstract: Vision foundation models (VFMs) such as DINO are pretrained for single-image representation, whereas remote sensing change detection requires reasoning

ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection

Model ReleasesDGX agent

arXiv:2608.09789v1 Announce Type: new Abstract: Industrial anomaly detection (IAD) requires identifying fine-grained deviations from normal visual patterns. Multimodal large language models (MLLMs) ca

Advantage-Guided Gate: Reshaping Open-Ended Reasoning for Vision-Based Spatial Intelligence

Model ReleasesDGX agent

arXiv:2608.07987v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential in complex spatial scene understanding and reasoning tasks. However, th

Adversarial Attacks on Deep OCR Systems

Model ReleasesDGX agent

arXiv:2608.07636v1 Announce Type: cross Abstract: Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at l

Adversarial Latent-State Training for Robust Policies in Partially Observable Domains

Model ReleasesDGX agent

arXiv:2603.07313v4 Announce Type: replace-cross Abstract: Robustness under latent distribution shift remains challenging in partially observable reinforcement learning. We formalize a focused setting

AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization

Model ReleasesDGX agent

arXiv:2608.07557v1 Announce Type: cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent minimalist en

AeroReformer2: Spoken-Query Referring Segmentation for Aerial Images

Model ReleasesDGX agent

arXiv:2608.08874v1 Announce Type: new Abstract: Spoken language offers a natural, hands-free interface for specifying an arbitrary target in dense remote-sensing imagery, yet existing referring remote

Agentic AI for Clustering, Relationship Discovery, and Semantic Trading in Prediction Markets

Model ReleasesDGX agent

arXiv:2512.02436v2 Announce Type: replace Abstract: Prediction markets allow users to trade on outcomes of real-world events, but are prone to fragmentation with overlapping questions, implicit equiva

Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data

Model ReleasesDGX agent

arXiv:2608.08859v1 Announce Type: cross Abstract: Wireless Body Area Networks (WBANs) generate multivariate physiological time series that are highly nonstationary and must often be processed under st

AgriField-40K: Adapting Vision Models to Agriculture With Efficient Continual Pretraining

Model ReleasesDGX agent

arXiv:2608.07984v1 Announce Type: new Abstract: Field-based agricultural computer vision is important for precision agriculture, yet it largely depends on expensive annotations and costly adaptation o

AI Evaluation Should Measure Verification Cost, Not Correctness Alone

Model ReleasesDGX agent

arXiv:2608.08709v1 Announce Type: new Abstract: The reliability of AI generative models is typically measured by output correctness, yet in practice it depends on the effort required to verify those o

An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography

Model ReleasesDGX agent

arXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency.

An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning

Model ReleasesDGX agent

arXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated b

Analysis and experiments of the dissipative Twistcar: direction reversal and asymptotic approximations

Model ReleasesDGX agent

arXiv:2506.19112v3 Announce Type: replace Abstract: Underactuated wheeled vehicles are commonly studied as nonholonomic systems with periodic actuation. Twistcar is a classical example inspired by a r

Anchor-Based AI Approach for Pre-Crash Object Detection Utilizing Micro-Doppler Signatures in Automotive Radar

Model ReleasesDGX agent

arXiv:2608.08701v1 Announce Type: new Abstract: Advanced automated driving presents significant potential to improve modern automotive safety systems, but it depends highly on the reliable activation

AndroidReality: How Far Are Mobile Agents from the Real World?

Model ReleasesDGX agent

arXiv:2608.07775v1 Announce Type: new Abstract: Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply in real-worl

← Previous
123456…369
Next →