AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,131 results
Model Releases

ADABORD: a novel AdaBoost approach for ordinal classification

DGX agent

arXiv:2607.21003v1 Announce Type: new Abstract: Ordinal Classification (OC) deals with classification tasks where the classes follow a natural order. Despite the progress in OC, many existing approach

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Adaptive Multi-Horizon Reinforcement Learning

DGX agent

arXiv:2607.20656v1 Announce Type: cross Abstract: Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks

DGX agent

arXiv:2607.21482v1 Announce Type: new Abstract: Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Agentic Designer: Progressive Multi-Agent Collaboration for Structure-Aware Interior Layout Generation

DGX agent

arXiv:2607.20866v1 Announce Type: new Abstract: Generating realistic interior furniture layouts that strictly adhere to architectural constraints (e.g., walls, doors, and windows) remains a fundamenta

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

AI Assistants Overassist

DGX agent

arXiv:2607.21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance from AI assis

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs

DGX agent

arXiv:2607.20498v1 Announce Type: new Abstract: Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve complex, long-h

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

An Analytically Trained Variational Surrogate for Quantum Phase Estimation on NISQ Hardware

DGX agent

arXiv:2607.20943v1 Announce Type: cross Abstract: Quantum Phase Estimation (QPE) is a foundational algorithm for molecular ground-state energy estimation, but its deep circuit requirements make direct

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

An LLM-Driven Workflow for Automated Process Control Strategy Generation and Tuning from Dynamic Process Models

DGX agent

arXiv:2607.21292v1 Announce Type: new Abstract: We present a structured large-language-model-driven workflow for automated multi-variable control design from dynamic process models. The workflow decom

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

DGX agent

arXiv:2607.20536v1 Announce Type: new Abstract: Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management

DGX agent

arXiv:2607.20764v1 Announce Type: new Abstract: We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard task-relevan

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

DGX agent

arXiv:2607.20596v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE fami

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

AREX: Towards a Recursively Self-Improving Agent for Deep Research

DGX agent

arXiv:2607.21461v1 Announce Type: new Abstract: Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candida

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models

DGX agent

arXiv:2607.20524v1 Announce Type: new Abstract: Mean cross-positional attention degradation is widely reported in transformer interpretability, yet whether it causally limits contextual retrieval rema

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Attribution Markets: A Fisher-Market Formulation for Fractional Credit Assignment Between Planned Tasks and Performed Actions

DGX agent

arXiv:2607.20694v1 Announce Type: new Abstract: Personal and organizational planning systems maintain two records that drift apart: what was planned (a task's effort budget) and what was done (a logge

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Automated Synthesis and Adversarial Validation of Executable Causal Research Pipelines

DGX agent

arXiv:2607.21173v1 Announce Type: new Abstract: While automated research systems promise to accelerate empirical analysis, they are prone to silent failures: instances in which analysis code executes

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Autonomous disproofs of the sum-product conjecture over mathbb R with GPT-5.5 Pro

DGX agent

arXiv:2607.20525v1 Announce Type: new Abstract: OpenAI's recent disproof of the Erdos unit distance conjecture marked a milestone for AI in mathematics. It also inspired another breakthrough: a human

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants

DGX agent

arXiv:2607.20488v1 Announce Type: new Abstract: Multi-agent LLM frameworks typically fix their team topology at boot time. When an individual agent becomes overloaded at runtime, for example by mixing

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

DGX agent

arXiv:2607.21588v1 Announce Type: new Abstract: Learning effective robot manipulation policies requires diverse, high-quality demonstrations, yet existing data pipelines are often difficult to scale b

model-releasesarxiv-cs-ro
24 Jul 2026
Model Releases

Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment

DGX agent

arXiv:2607.20476v1 Announce Type: new Abstract: We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data. Testing 60 scenarios across three c

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Benchmarking the Personalization Capabilities of Large Language Models

DGX agent

arXiv:2607.20471v1 Announce Type: new Abstract: Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has a long tradit

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Benchmarking Unlearning for Vision Transformers

DGX agent

arXiv:2602.20114v2 Announce Type: replace-cross Abstract: Machine unlearning (MU) refers to the post-training capability to remove (the influence of) training examples that are incorrect, biased, or l

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs

DGX agent

arXiv:2607.20479v1 Announce Type: new Abstract: Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Break Through the Compression Bottleneck: From Theory to Practice

DGX agent

arXiv:2607.20434v1 Announce Type: cross Abstract: As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memory overhead.

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

CAMeR: Keyword-Gated Hybrid Activation for Adaptive Memory Retention in LLM Agents

DGX agent

arXiv:2607.20458v1 Announce Type: cross Abstract: Large language model (LLM) agents operating over extended dialogues accumulate vast amounts of information, yet existing memory systems either retain

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits

DGX agent

arXiv:2607.20518v1 Announce Type: new Abstract: AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Existing benchma

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

DGX agent

arXiv:2607.21340v1 Announce Type: new Abstract: In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft, but whether the draft is bankable: defens

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Cardinality-Decomposed Loss: Matching Training Objectives to Relation Structure in Heterogeneous Recommendation Graphs

DGX agent

arXiv:2607.20737v1 Announce Type: new Abstract: Graph Neural Networks trained on heterogenous bipartite graphs form a common basis in recommendation systems. These graphs often express relations that

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Case study: solving P-99 with LPTP and an LLM

DGX agent

arXiv:2607.21196v1 Announce Type: cross Abstract: Ninety-Nine Prolog Problems (P-99) is a famous set of Prolog exercises. We solved the first thirty three just by prompting an LLM (Large Language Mode

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Chronofy: A Temporal-Logical Decay Architecture for Information Validity in Time-Aware Retrieval-Augmented Generation

DGX agent

arXiv:2607.20560v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems retrieve and integrate external knowledge to ground large language model (LLM) outputs. However, current RA

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Classifier-pruned Bayesian optimization for particle accelerator tuning: Exploring temporally structured manifold of 6D beam phase space

DGX agent

arXiv:2412.01748v2 Announce Type: replace Abstract: Complex dynamical systems, such as particle accelerators, often require intricate and time-consuming tuning procedures to achieve optimal performanc

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

ConfidenceBench: Evaluating Confidence Calibration in Large Language Models

DGX agent

arXiv:2607.20526v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, accuracy alone i

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA

DGX agent

arXiv:2607.21155v1 Announce Type: cross Abstract: Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiri

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

CT-Merging: Consensus Directions and Task-Level Scaling for LoRA Adapter Merging

DGX agent

arXiv:2607.20561v1 Announce Type: cross Abstract: LoRA adapters provide an efficient way to specialize a pretrained model for many downstream tasks, but deploying one adapter per task requires adapter

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

CultureTalk-ID: A Multi-Task Dialogue Benchmark for Cultural Commonsense in Indonesian Local Languages

DGX agent

arXiv:2607.21016v1 Announce Type: new Abstract: Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs on short and isolated prompts, stripping aw

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Cycle-Consistent and Uncertainty-Aware Neural Surrogates for Tokamak Edge Plasmas

DGX agent

arXiv:2607.21407v1 Announce Type: cross Abstract: The boundary and divertor plasma govern how a tokamak exhausts power and particles, setting heat fluxes, target conditions, and the onset of detachmen

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

DGX agent

arXiv:2607.20465v1 Announce Type: cross Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining

DGX agent

arXiv:2603.11838v2 Announce Type: replace Abstract: Large language models pretrained on internet-scale data risk lookahead bias in forecasting tasks, as they may have already seen the true outcome dur

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Deblurring in the Wild: A Real-World Image Deblurring Dataset from Smartphone High-Speed Videos

DGX agent

arXiv:2506.19445v4 Announce Type: cross Abstract: We introduce the largest real-world image deblurring dataset constructed from smartphone slow-motion videos. Using 240 frames captured over one second

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Demonstrating GenDB: Instance-Optimized and Customized Query Processing Code Generation via LLM Agents

DGX agent

arXiv:2607.20630v1 Announce Type: cross Abstract: Traditional query processing engines require continuous development and extensions to support new techniques and user requirements, and in some cases,

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Detecting Neural Network Failures through Spectral Analysis of Internal Activations

DGX agent

arXiv:2607.20590v1 Announce Type: cross Abstract: Neural network misclassifications exhibit characteristic spectral instability in internal activations that is invisible at the output layer. This phen

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making

DGX agent

arXiv:2607.20491v1 Announce Type: new Abstract: Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time. We i

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

DINO-VPT: Hierarchical Visual Prompt Tuning for Joint Physical-Digital Face Anti-Spoofing

DGX agent

arXiv:2607.20900v1 Announce Type: new Abstract: With the increasing diversity of spoofing attacks, there is a growing demand for unified Face Anti-Spoofing (FAS) models capable of detecting both physi

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

Do Active SAE Feature Planes Carry More Holonomy? A Preregistered Reversal in Gemma

DGX agent

arXiv:2607.20522v1 Announce Type: new Abstract: This paper tests whether holonomy concentrates on active sparse-autoencoder (SAE) feature planes in Gemma 2 2B, a concrete operationalization of the bro

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Do Pathology Vision-Language Models Truly See Pathology?

DGX agent

arXiv:2607.21065v1 Announce Type: new Abstract: Pathology vision-language models (VLMs) have recently progressed rapidly and are commonly evaluated by answer accuracy on pathology VQA benchmarks. Howe

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

Domyn-Small: A European 10B Reasoning Language Model

DGX agent

arXiv:2607.20448v1 Announce Type: new Abstract: We introduce Domyn-Small, a 10-billion-parameter open-weight reasoning language model released under the MIT license. Domyn-Small is the product of an i

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

DGX agent

arXiv:2607.21540v1 Announce Type: new Abstract: We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Drive As You Like: Multi-Head Diffusion with Reinforcement Learning for Personalized Driving

DGX agent

arXiv:2508.16947v2 Announce Type: replace-cross Abstract: Despite significant progress, imitation learning-based autonomous driving planners remain largely restricted to reproducing high-frequency bia

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention

DGX agent

arXiv:2607.20457v1 Announce Type: cross Abstract: Inference with large language models (LLMs) on long sequences is computationally expensive due to the quadratic complexity of self-attention. Distribu

model-releasesarxiv-cs-ai
24 Jul 2026
← Previous
1…5859606162…357
Next →