AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
All
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,512 results
Model Releases

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

DGX agent

arXiv:2607.24063v1 Announce Type: new Abstract: On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a system is tow

model-releasesarxiv-cs-ai
28 Jul 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

The Few-shot Dilemma: Over-prompting Large Language Models

DGX agent

arXiv:2509.13196v2 Announce Type: replace Abstract: Over-prompting, a phenomenon where excessive examples in prompts lead to diminished performance in Large Language Models (LLMs), challenges the conv

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

The Gate Always Closes: On Injecting Auxiliary Signals into Frozen Vision-Language Models

DGX agent

arXiv:2607.23335v1 Announce Type: new Abstract: Auxiliary signal pathways in VLMs are routinely fitted with learnable gates so the optimiser can decide how much of the signal to admit. We find that th

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

DGX agent

arXiv:2607.24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

The Kimi K3 architecture figure for yesterday's big open-weight model release, along with some observations and thoughts. 1. Yes, it looks r…

DGX agent

The Kimi K3 architecture figure for yesterday's big open-weight model release, along with some observations and thoughts. 1. Yes, it looks relatively complicated, but it's essentially a scaled-up prod

model-releasessebastian-raschka--x
28 Jul 2026
Model Releases

The Label Complexity of Class-Conditional Coverage under Distribution Shift

DGX agent

arXiv:2607.18088v2 Announce Type: replace-cross Abstract: Conformal prediction certifies that a classifier's prediction sets cover the truth, and that certificate is marginal. Many recognition benchma

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation

DGX agent

arXiv:2607.22585v1 Announce Type: new Abstract: Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools,

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages

DGX agent

arXiv:2607.24276v1 Announce Type: cross Abstract: Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are tr

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

ThinkingCap-Qwen3.6-27B warrants a look

DGX agent

It has only been two days since I move 100% from Qwen3.5-27B F16 to ThinkingCap-Qwen3.6-27B F16. Where I was getting tps in 30-40 range (depending on the size of the context), I am definitely getting

model-releasesr-localllama
28 Jul 2026
Model Releases

Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models

DGX agent

arXiv:2607.23054v1 Announce Type: cross Abstract: Multi-head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key-value pairs through a shared low-rank bottleneck (cKV), achieving 81% KV-

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding

DGX agent

arXiv:2607.23951v1 Announce Type: new Abstract: Video temporal grounding (VTG) aims to localize the continuous video interval described by a natural-language query. However, current VLM-based methods

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

Tines seeks to tame ‘wild code’ AI sprawl

DGX agent

Workflow automation company Tines Security Services Ltd. today launched Tines 3B, an artificial intelligence-native platform for building, running and governing enterprise workflows, applications and

model-releasessiliconangle
28 Jul 2026
Model Releases

TLA^{+}-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA+ Specification Generation

DGX agent

arXiv:2607.23425v1 Announce Type: cross Abstract: Large language models increasingly write TLA^{+} formal specifications from natural-language descriptions, but progress is hard to measure: existing r

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

TLRNet: Estimating Individual Treatment Effect based on Local Information and Single Learner Structure

DGX agent

arXiv:2607.22762v1 Announce Type: cross Abstract: Causal inference has become a central issue across various fields, including computer science, statistics, economics, education, healthcare, and medic

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations

DGX agent

arXiv:2607.22610v1 Announce Type: new Abstract: When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those dependencies

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

TokenMem: Faithful Knowledge Injection for Frozen LLMs

DGX agent

arXiv:2607.22625v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge, but suffers from knowledge conflicts: when retrieved

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records

DGX agent

arXiv:2607.22954v1 Announce Type: new Abstract: Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world d

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

Towards simultaneous decoding of kinetic and kinematic movement parameters during grasp and lift task by noninvasive brain imaging

DGX agent

arXiv:2607.24081v1 Announce Type: cross Abstract: Brain-machine interfaces (BMIs) can assist individuals with limited mobility, such as stroke survivors or amputees. One of the key challenges in devel

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Transfer Learning Architectures for Scalable Multi-Fidelity Bayesian Optimization

DGX agent

arXiv:2607.23404v1 Announce Type: new Abstract: Self-driving laboratories increasingly rely on multi-fidelity Bayesian optimization (MFBO) to balance cheap, approximate evaluations against scarce, exp

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

TRE: Training-Free Hallucination Detection for Diffusion Language Models

DGX agent

arXiv:2607.22661v1 Announce Type: new Abstract: Diffusion large language models (D-LLMs) have recently gained increasing attention, yet their reliability is significantly hindered by the hallucination

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

DGX agent

arXiv:2507.21134v2 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in high-risk domains such as law, finance, and medicine, systematically evaluating their d

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation

DGX agent

arXiv:2607.23838v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base at query t

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

TriSP: Tri-Signal Structured Pruning for Large Language Models

DGX agent

arXiv:2607.22587v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance across diverse tasks but their deployment is constrained by the memory and compute cost of their

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Trustworthy Medical Segmentation: Uncertainty-Aware U-Net Evaluation Under Clinical Image Degradation

DGX agent

arXiv:2607.22727v1 Announce Type: new Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

Try Once, Then Optimal: De-Redundified Procedure Memory for Cross-Episode Exploration Amortization

DGX agent

arXiv:2607.23702v1 Announce Type: new Abstract: Manipulating objects with hidden internal state, such as a latched microwave, forces a robot to probe before it can act. Yet a robot that has solved an

model-releasesarxiv-cs-ro
28 Jul 2026
Model Releases

Two Regimes of Chain-of-Thought Unfaithfulness: Behavioral Detection Fails Where Models Are Wrong

DGX agent

arXiv:2607.23458v1 Announce Type: new Abstract: Chain-of-thought (CoT) explanations support oversight only if they are faithful: the stated reasoning must actually produce the answer. Auditing black-b

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

Two-Timescale Hierarchical Reinforcement Learning for Resilient Operations

DGX agent

arXiv:2607.23434v1 Announce Type: cross Abstract: Unexpected shocks recur in global operations, requiring decision rules that adapt as market and operating conditions change. Many operational systems

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents

DGX agent

arXiv:2508.00288v5 Announce Type: replace-cross Abstract: Aerial navigation is a fundamental yet underexplored capability in embodied intelligence, enabling agents to operate in large-scale, unstructu

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

Understanding Machine Unlearning Through the Lens of Mode Connectivity

DGX agent

arXiv:2607.23970v1 Announce Type: cross Abstract: Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress, the loss la

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Understanding Tone-Dependent Inference Cost in Large Language Models

DGX agent

arXiv:2607.23915v1 Announce Type: cross Abstract: We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were perf

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training

DGX agent

arXiv:2512.24125v3 Announce Type: replace-cross Abstract: General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action executio

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective

DGX agent

arXiv:2607.24507v1 Announce Type: cross Abstract: Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform-noise diff

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Unsupervised Graph Representation Learning with Complementary View Alignment

DGX agent

arXiv:2607.24338v1 Announce Type: new Abstract: Unsupervised graph representation learning aims to derive meaningful node embeddings by capturing both structural and attribute information without rely

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

UP-Fuse: Uncertainty-guided LiDAR-Camera Fusion for 3D Panoptic Segmentation

DGX agent

arXiv:2602.19349v2 Announce Type: replace-cross Abstract: LiDAR-camera fusion enhances 3D panoptic segmentation by leveraging camera images to complement sparse LiDAR scans, but it also introduces a c

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Update your chat template for dsv4 if you're using llama.cpp

DGX agent

Following some recent commits in llama.cpp, preserve_thinking behavior for chat templates included in older DSV4 ggufs got broken. This makes the model pretty dumb in a coding agent context. Adding kw

model-releasesr-localllama
28 Jul 2026
Model Releases

Using Reinforcement Learning to Optimize the Global and Local Crossing Number

DGX agent

arXiv:2509.06108v3 Announce Type: replace-cross Abstract: Graph drawing concerns the algorithmic visualization of graphs. A good drawing of a graph is easy to read and facilitates solving tasks on the

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

Variational Quantum Conditional Boltzmann Machines for Time-Series Forecasting: Architectures, Symmetric Hyperparameter Evaluation, and a Nonlinear Benchmark

DGX agent

arXiv:2607.24065v1 Announce Type: cross Abstract: In this study, we developed and evaluated four conditional energy-based forecasting architectures: a classical Gaussian-Bernoulli CRBM, a hybrid quant

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

Verbalized Particle Posterior: Bayesian Inference over Natural Language Hypotheses

DGX agent

arXiv:2607.22961v1 Announce Type: cross Abstract: Verbalized Machine Learning (VML) parameterizes a model as a natural-language prompt that an LLM evaluates as f(x; theta). The framework is interpreta

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model

DGX agent

arXiv:2607.22704v1 Announce Type: cross Abstract: Nuclear fusion has made significant progress in recent years and is expected to become one of the most important pathways to addressing global energy

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs

DGX agent

arXiv:2512.22342v5 Announce Type: replace Abstract: In most existing embodied navigation tasks, instructions are well-defined and unambiguous, such as instruction following and object searching. Under

model-releasesarxiv-cs-ro
28 Jul 2026
Model Releases

VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing

DGX agent

arXiv:2607.22632v1 Announce Type: new Abstract: The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editing plans. Howe

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Weighted Low-Rank Matrix Approximation: Acceleration and Applications

DGX agent

arXiv:2109.11057v2 Announce Type: replace-cross Abstract: Weighted low-rank matrix approximation (WLRMA) generalizes classical low-rank approximation and matrix completion by allowing arbitrary elemen

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

What would it take for the frontier labs to open the weights of their old, deprecated proprietary models?

DGX agent

Anyone thought about this? What do you think needs to happen for them to release the old weights? I’d love to see models like Gemini-2.5, OAI o3, 4o, 4.1 being open one day. In Oct 2025 Scam Altman sa

model-releasesr-localllama
28 Jul 2026
Model Releases

When Less Is More: A Controlled Benchmark of Lightweight CNNs for Satellite Land-Cover Segmentation on DeepGlobe

DGX agent

arXiv:2607.23024v1 Announce Type: new Abstract: High-resolution satellite imagery is the backbone of good land-cover classification, and without that, environmental monitoring, urban planning, and sus

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

DGX agent

arXiv:2607.24392v1 Announce Type: cross Abstract: Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. W

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents

DGX agent

arXiv:2607.24077v1 Announce Type: new Abstract: Optical Character Recognition (OCR) is a key component in the digitization of historical archives. Recently, Vision-Language Models (VLMs) have emerged

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

Which Models Perform Better in Inheritance Reasoning?

DGX agent

arXiv:2606.13751v4 Announce Type: replace Abstract: This paper presents the participation of team PSL in the QIAS 2026 Shared Task on Arabic Islamic inheritance reasoning. The task evaluates the abili

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It

DGX agent

arXiv:2607.23893v1 Announce Type: cross Abstract: Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the questio

model-releasesarxiv-cs-cl
28 Jul 2026
← Previous
1…8081828384…469
Next →