AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,545 results
Model Releases

ODeform: Learning Continuous 4D Motion for Shape Deformation with Neural ODEs

DGX agent

arXiv:2607.20670v1 Announce Type: new Abstract: Modeling continuous object deformation is important for many computer vision and robotics tasks, such as manipulation and simulation. Existing approache

model-releasesarxiv-cs-cv
24 Jul 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and quality, Opus 5 scores 63.6% with a 69.6% p…

DGX agent

On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and quality, Opus 5 scores 63.6% with a 69.6% pass rate on Extended — approaching Fable 5 at half the cost.

model-releasescognition-ai--x
24 Jul 2026
Model Releases

One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies

DGX agent

arXiv:2607.21143v1 Announce Type: cross Abstract: Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to a

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Open Knowledge format v0.2 tackles agentic trust

DGX agent

When we introduced the Open Knowledge Format (OKF) in June 2026, we asserted that the context that agents need (table schemas, metric definitions, runbooks) should live in a format, not in a proprieta

model-releasesgoogle-cloud-ai
24 Jul 2026
Model Releases

Open Source Tax Engine outperforming fable 5 and gpt sol

DGX agent

This is an open source and free tax engine which scored 96% on TaxCalcBench [highest ever recorded score till date] surpassing fable 5 and sol with just sonnet 5. The only 2 cases where it missed, it

model-releasesr-ollama
24 Jul 2026
Model Releases

OpenForgeRL: Train Harness-native Agents in Any Environment

DGX agent

arXiv:2607.21557v1 Announce Type: new Abstract: Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to e

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Optimizing an Ollama (Qwen:2.5) AI Agent: Fixing Search Aggregation, Context Bleed, and Query Extraction

DGX agent

I am building a domain-specific AI agent powered by Ollama (using the qwen:2.5 model). For data retrieval, the agent utilizes multiple search APIs: DuckDuckGo Search (DDGS), Tavily, Serper, and Google

model-releasesr-ollama
24 Jul 2026
Model Releases

Opus 5 improves coding, reasoning efficiency, and prompt-cache-friendly tool use, and is priced at 5/1M input tokens and 25/1M output, same as Opus 4.8 (David Gewirtz/ZDNET)

DGX agent

David Gewirtz / ZDNET: Opus 5 improves coding, reasoning efficiency, and prompt-cache-friendly tool use, and is priced at 5/1M input tokens and 25/1M output, same as Opus 4.8 — ZDNET's key takeaways —

model-releasestechmeme
24 Jul 2026
Model Releases

Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most excitin…

DGX agent

Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt inject

model-releasesboris-cherny--x
24 Jul 2026
Model Releases

Opus 5 now available in Hermes Agent

DGX agent

Claude Opus 5 is now released in the Hermes Agent, a product of Nous Research and Teknium. Users can access the model through multiple gateways, including the Nous Portal, OpenRouter, and Anthropic Di

model-releasesnous-research--x
24 Jul 2026
Model Releases

pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development

DGX agent

arXiv:2607.21268v1 Announce Type: cross Abstract: In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable co

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

[Paper] Statistically-Lossless Quantization of Large Language Models

DGX agent

Model quantization has become essential for efficient large language model deployment, yet existing approaches involve clear trade-offs: methods such as GPTQ and AWQ achieve practical compression but

model-releasesr-localllama
24 Jul 2026
Model Releases

People are using Minecraft farms as AI agent benchmarks

DGX agent

Someone modelled sugarcane farming as an integer program. See, sugarcane only grows next to water. Water costs one tile and can feed at most four cane tiles. The layout therefore becomes a coverage pr

model-releasesr-chatgpt
24 Jul 2026
Model Releases

PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

DGX agent

arXiv:2607.20482v1 Announce Type: new Abstract: Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspeci

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

PhantomFill: When the Form Demands an Answer, Language Models Invent One

DGX agent

arXiv:2607.20492v1 Announce Type: cross Abstract: Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses

DGX agent

arXiv:2603.13026v2 Announce Type: replace Abstract: Prompt injection poses serious security risks to real-world LLM applications, particularly autonomous agents. Although many defenses have been propo

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks

DGX agent

arXiv:2607.20864v1 Announce Type: cross Abstract: Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single ans

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Preference Tuning as Spectral Update Reorganization

DGX agent

arXiv:2607.20438v1 Announce Type: cross Abstract: Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains largely opa

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning

DGX agent

arXiv:2607.21022v1 Announce Type: new Abstract: Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing t

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

Profiling Lightweight Large Language Models

DGX agent

arXiv:2607.20806v1 Announce Type: new Abstract: Lightweight large language models (LLMs) are increasingly being deployed locally on personal computers and are expected to play a growing role in resour

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

PromptPack: Scaling LLM Annotation Agents for Online Recommendation

DGX agent

arXiv:2607.20528v1 Announce Type: new Abstract: Online recommendation platforms increasingly use Large Language Models (LLMs) to extract structured features from ad creatives. While deploying a single

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

DGX agent

arXiv:2607.21063v1 Announce Type: new Abstract: Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is as

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

RE-AD: Real-Time Requirement Adherence for Data Labeling

DGX agent

arXiv:2607.20455v1 Announce Type: cross Abstract: Human-annotated data remains fundamental to training frontier Large Language Models (LLMs). However, crowd-sourced annotations often suffer from quali

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

DGX agent

arXiv:2607.20791v1 Announce Type: new Abstract: High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances in truncation-based sampling techniques hav

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

REGARD: Regional Affective Differences in Large Language Models

DGX agent

arXiv:2607.20722v1 Announce Type: new Abstract: Large language models trained and aligned within different linguistic and regional ecosystems may frame the same political, cultural, and geopolitical e

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Relative Value Learning

DGX agent

arXiv:2607.21120v1 Announce Type: cross Abstract: In reinforcement learning, critics typically estimate absolute state values V(s), estimating how good a particular situation is in isolation. However,

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Replit shipped a lot this month. Talk to Agent with your voice, build from Claude or Slack, and never re-explain your stack to Agent again. …

DGX agent

Replit released several major features this month, including a voice‑enabled Agent that lets users talk directly to the platform. Users can now build from Claude or Slack, and the updated Agent rememb

model-releasesreplit--x
24 Jul 2026
Model Releases

RUMBA: Russian User Memory Benchmark

DGX agent

arXiv:2607.21447v1 Announce Type: cross Abstract: The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely on aggregate

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Rushes: A Human Preference Dataset for Pluralistic Alignment

DGX agent

arXiv:2607.20767v1 Announce Type: new Abstract: We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collect

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

DGX agent

arXiv:2607.21518v1 Announce Type: new Abstract: Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction.

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Scaling Closed-Loop Feature Channel Configuration with LLMs

DGX agent

arXiv:2607.20516v1 Announce Type: cross Abstract: Promising initial results in closed-loop large-language-model-based channel-configuration search demonstrated that neural-network widths can be optimi

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Scaling Interpretable Transformers with Parity Bottleneck Layers

DGX agent

arXiv:2607.20652v1 Announce Type: cross Abstract: Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Spa

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Scene Parameter Saliency via Differentiable Light Transport

DGX agent

arXiv:2607.21562v1 Announce Type: new Abstract: Gradient-based saliency methods reveal which input features most influence a neural network's output, and are a standard tool for model interpretability

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

DGX agent

arXiv:2607.20926v1 Announce Type: new Abstract: Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily em

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents

DGX agent

arXiv:2602.10226v2 Announce Type: replace-cross Abstract: Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hyper

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Semi-Supervised Text-Attributed Graph Distillation

DGX agent

arXiv:2607.20477v1 Announce Type: new Abstract: {em Text-Attributed Graphs} (TAGs) have emerged as an expressive data model for integrating graph topology with rich textual semantics. Existing represe

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

SenCos-GEM: SENet-Calibrated and Law-of-Cosines-Constrained Geometry-Enhanced Molecular Representation for Property Prediction

DGX agent

arXiv:2607.20551v1 Announce Type: cross Abstract: Effective molecular representation learning is crucial for accurate molecular property prediction. Recently, numerous self-supervised learning (SSL) a

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

SESaMo: Symmetry-Enforcing Stochastic Modulation for Normalizing Flows

DGX agent

arXiv:2505.19619v3 Announce Type: replace Abstract: Deep generative models have recently garnered significant attention across various fields, from physics to chemistry, where sampling from unnormaliz

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

DGX agent

arXiv:2607.21072v1 Announce Type: new Abstract: Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks a

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

DGX agent

arXiv:2607.09999v2 Announce Type: replace Abstract: We show that post-training quantization can silently alter how large language models reason even when task accuracy is preserved. Using a six-catego

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents

DGX agent

arXiv:2607.15557v4 Announce Type: replace Abstract: Agent skills, SKILL files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Pub

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

DGX agent

arXiv:2607.20548v1 Announce Type: cross Abstract: Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges hav

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification

DGX agent

arXiv:2607.20475v1 Announce Type: new Abstract: Sampling in LLM inference comprises a combinatorial set of logit processing, token selection, and verification operations for speculative decoding. Howe

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Source-Prior-Driven Selective Adaptation for Efficient Diffusion Model Finetuning

DGX agent

arXiv:2607.20913v1 Announce Type: new Abstract: Fine-tuning large diffusion models for new domains or styles involves a trade-off: improving target-specific generation often degrades the pretrained mo

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Spectral functions in Minkowski quantum electrodynamics from neural reconstruction

DGX agent

arXiv:2510.24728v2 Announce Type: replace-cross Abstract: We study neural reconstructions of quenched rainbow quantum electrodynamics (QED) Dyson--Schwinger benchmarks in Minkowski-related kinematics.

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Spectral-Spatial Synergistic Guided Network for Hyperspectral Salient Object Detection

DGX agent

arXiv:2607.21032v1 Announce Type: new Abstract: Hyperspectral salient object detection aims to identify visually salient regions from hyperspectral images. Existing methods often fail because they fun

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

StabilityBench: Benchmarking Instability in LLMs

DGX agent

arXiv:2607.20558v1 Announce Type: cross Abstract: AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poor

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Statistical Inference for Generative Model Comparison

DGX agent

arXiv:2501.18897v4 Announce Type: replace-cross Abstract: Generative models have achieved remarkable success across a range of applications, yet their evaluation still lacks principled uncertainty qua

model-releasesarxiv-cs-lg
24 Jul 2026
← Previous
1…8990919293…470
Next →