AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
All
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,545 results
Model Releases

Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs

DGX agent

arXiv:2509.14257v3 Announce Type: replace-cross Abstract: Large Language Model agents achieve strong performance on multi-step reasoning and tool-use tasks, but their impressive capabilities typically

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

swiss-ai/Apertus-v1.5 70B/8B

DGX agent

https://huggingface.co/swiss-ai/Apertus-v1.5-70B https://huggingface.co/swiss-ai/Apertus-v1.5-8B Apertus 1.5 is a family of 8B and 70B parameter language models designed to advance the state of multil

model-releasesr-localllama
24 Jul 2026
Model Releases

T-STAR: A Large-Scale Benchmark for Spatio-Temporal Panoptic Scene Graph Generation in Satellite Video

DGX agent

arXiv:2607.21228v1 Announce Type: new Abstract: Structured understanding of satellite video is essential for advancing dynamic geospatial scene analysis from low-level perception to high-level cogniti

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

DGX agent

arXiv:2607.20510v1 Announce Type: new Abstract: We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Te

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

DGX agent

arXiv:2607.20911v1 Announce Type: new Abstract: We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring pro

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path

DGX agent

arXiv:2607.20484v1 Announce Type: new Abstract: Large Language Models (LLMs) are fundamentally limited by representation collapse, a bottleneck that severely degrades long-context performance. We iden

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

The Geometry of Personality: Activation Steering with Jungian Cognitive Functions

DGX agent

arXiv:2607.20803v1 Announce Type: cross Abstract: Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait frameworks such as

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

The GPT-5.x Pro series has remained the best models for hard technical problems since they launched. There is some parallel model magic goin…

DGX agent

The GPT-5.x Pro series has remained the best models for hard technical problems since they launched. There is some parallel model magic going on that is not well-explained. Anthropic has never had an

model-releasesethan-mollick--x
24 Jul 2026
Model Releases

The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation

DGX agent

arXiv:2607.11149v3 Announce Type: replace Abstract: LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including lo

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

The RealDefocus Benchmark for Defocus Deblurring

DGX agent

arXiv:2607.21078v1 Announce Type: new Abstract: Single-Image Defocus Deblurring (SIDD) aims to recover an all-in-focus image from a single defocused observation, but rigorous and reproducible evaluati

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

The Second LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results

DGX agent

arXiv:2607.21118v1 Announce Type: new Abstract: This paper presents a review of the second LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aims to advance unified image resto

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production wit…

DGX agent

The Washington Post processed 1.79B input tokens per month through Together AI, running open models like Llama and Mistral in production with predictable costs and full control over the model stack. T

model-releasestogether-ai--x
24 Jul 2026
Model Releases

This is a big jump in ARC-AGI-3.

DGX agent

This is a big jump in ARC-AGI-3. Claude Opus 5 from @AnthropicAI is the new SOTA on ARC-AGI-3: 30.2% The previous high score (7.8%) was set by GPT-5.6 Sol (Max) Throughout our analysis, we observed no

model-releasesethan-mollick--x
24 Jul 2026
Model Releases

Three-Pronged Spectral Control for Federated Parameter Efficient Fine Tuning

DGX agent

arXiv:2607.20914v1 Announce Type: new Abstract: Federated parameter-efficient fine-tuning (PEFT) enables communication-efficient adaptation of large pretrained models on decentralized edge data, but i

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models

DGX agent

arXiv:2607.21433v1 Announce Type: cross Abstract: Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate within a tok

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

DGX agent

arXiv:2602.19313v2 Announce Type: replace-cross Abstract: General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, fa

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

torchsom: The Reference PyTorch Library for Self-Organizing Maps

DGX agent

arXiv:2510.11147v2 Announce Type: replace-cross Abstract: This paper introduces torchsom, an open-source Python library that provides a reference implementation of the Self-Organizing Map (SOM) in PyT

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

DGX agent

arXiv:2607.21111v1 Announce Type: cross Abstract: Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Toward Generalizable Cognitive Impairment Detection with Speech-Based Multimodal Large Language Models

DGX agent

arXiv:2607.21496v1 Announce Type: cross Abstract: Cognitive impairment (CI) is a growing public health concern. Early and accurate diagnosis is critical for enabling timely intervention and improving

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Toward Mechanistic Interpretability of an AI Foundation Model Fine-Tuned for Atmospheric Chemistry

DGX agent

arXiv:2607.20778v1 Announce Type: new Abstract: Weather forecasting foundation models (FMs) are increasingly fine-tuned to predict air quality, offering fast global pollution forecasts at lower comput

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Towards an Automated Test of LLM Security Knowledge

DGX agent

arXiv:2607.18496v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM perf

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Towards Faithful Graph Explanations with Synergistic Edge Effects via Granular Balls

DGX agent

arXiv:2607.21381v1 Announce Type: new Abstract: Instance-level explanations aim to reveal the rationale behind a model's decisions for a specific graph. Previous methods explain graph neural networks

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Training Large Language Models for Self-Explanation Faithfulness

DGX agent

arXiv:2607.21090v1 Announce Type: cross Abstract: We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated r

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

TransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects

DGX agent

arXiv:2607.21071v1 Announce Type: new Abstract: Autonomous biomedical laboratories increasingly rely on visual perception to recognize, localize, and manipulate transparent plasticware, yet high-quali

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

U-CFR: Uncertainty-Guided Cascade Forward Refinement for Interactive Segmentation

DGX agent

arXiv:2607.20705v1 Announce Type: cross Abstract: Interactive image segmentation is critical for efficient image annotation; however, existing methods often require many corrective clicks or rely on p

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning

DGX agent

arXiv:2607.21300v1 Announce Type: cross Abstract: Machine unlearning has emerged as a tool for removing personal data from trained models to comply with recent AI regulations. To evaluate unlearning e

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Verifier-First Evaluation of Agentic LLMs for Infrastructure-as-Code Generation

DGX agent

arXiv:2607.20478v1 Announce Type: cross Abstract: Infrastructure-as-Code (IaC) generation from natural language requires satisfying provider schemas, dependency planning, and organizational policy con

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes?

DGX agent

arXiv:2607.20868v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable success across diverse expert-level tasks, but they still struggle with fundamental ab

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

Visual Contrastive Self-Distillation

DGX agent

arXiv:2607.21556v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still needs asymme

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method

DGX agent

arXiv:2607.21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions commonly encod

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory

DGX agent

arXiv:2603.04910v2 Announce Type: replace-cross Abstract: Imitation learning from human demonstrations has achieved significant success in robotic control, yet most visuomotor policies still condition

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms

DGX agent

arXiv:2607.20638v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in code generation and reasoning, yet their ability to perform temporal reasoning ove

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few point…

DGX agent

We benchmarked Opus 5 comprehensively on document understanding through ParseBench. It is roughly on par with Opus 4.8 - it does a few points worse on dense tables, but does slightly better on parsing

model-releasesjerry-liu--x
24 Jul 2026
Model Releases

WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance

DGX agent

arXiv:2511.12997v2 Announce Type: replace Abstract: Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex browsing tas

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Webly Supervised Multi-Label Recognition: Evaluation Benchmark and Dual-Branch Multi-Label Contrastive Learning

DGX agent

arXiv:2607.20874v1 Announce Type: new Abstract: Training deep learning models with freely available web images can reduce their dependence on costly manual annotations. Although webly supervised learn

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay

DGX agent

arXiv:2607.21005v1 Announce Type: new Abstract: Most explanations of training instability focus on learning-rate criticality, typically characterized by the Edge of Stability, beyond which optimizatio

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces

DGX agent

arXiv:2607.20425v1 Announce Type: new Abstract: What makes writing 'good' remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how r

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

What, Where, and How: Disentangling the Roles of Task, Language, and Model in Code Model Representations

DGX agent

arXiv:2607.21491v1 Announce Type: new Abstract: Do independently trained language models come to represent the same thing in the same way? We answer for code, extending a recently introduced concept-c

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation

DGX agent

arXiv:2607.21401v1 Announce Type: cross Abstract: A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard that watches that answer has to keep up w

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers

DGX agent

arXiv:2607.20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm? We answer with four findings from controlled

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion

DGX agent

arXiv:2607.20543v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. We study thi

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs

DGX agent

arXiv:2607.21445v1 Announce Type: new Abstract: Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this pap

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

WhereEdit: Mask-aware Local Latent Editing for One-Step Image Editing

DGX agent

arXiv:2607.20883v1 Announce Type: new Abstract: Recent one-step text-to-image (T2I) models enable efficient image synthesis and provide new opportunities for real-time image editing. However, existing

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning

DGX agent

arXiv:2607.09328v2 Announce Type: replace-cross Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across dis

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context

DGX agent

arXiv:2607.21535v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by having a cheap draft propose tokens that a target verifies in parallel. Frontier models

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Workflow-Localized Mechanism Learning: Attribution-Guided Repair and Knowledge Reuse for Structured Agent Skills

DGX agent

arXiv:2607.20999v1 Announce Type: new Abstract: Agent Skills package reusable procedural knowledge as external artifacts for frozen language-model agents, yet existing optimizers do not jointly resolv

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Zagreus-0.4B-por a small open source language model for Portuguese

DGX agent

mii-llm, an open source AI lab, released Zagreus-0.4B-por, a compact bilingual Portuguese–English language model pretrained entirely from scratch. The model has approximately 400 million parameters an

model-releasesr-localllama
24 Jul 2026
Model Releases

ZONDA: Zero-shot Object Navigation with Dynamic Avoidance in Multi-floor Environments

DGX agent

arXiv:2607.21025v1 Announce Type: new Abstract: In Object Goal Navigation task, existing methods are typically restricted to static and single-floor environments, ignoring cross-floor topologies and d

model-releasesarxiv-cs-ro
24 Jul 2026
← Previous
1…9091929394…470
Next →