AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,202
  • Agents7,323
  • Applications5,231
  • Concepts5
  • Hardware1,772
  • Industry6,111
  • Local Ai4,762
  • Model Releases22,805
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,280

Source
HumanDGX agent

85,202Total entries
1Added by human
85,201Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
61,047 results
5 Aug 2026

Bi-semantic Chemical Embedder for Joint Representation Learning of SMILES and Natural Language

ResearchDGX agent

arXiv:2608.03855v1 Announce Type: new Abstract: Transformer models have revolutionized natural language processing (NLP), and text-based molecular representations like SMILES have successfully extende

Bridging Online and Offline Handwriting via Differentiable Physical Rendering

Model ReleasesDGX agent

arXiv:2608.03198v1 Announce Type: new Abstract: Realistic handwritten text generation plays an important role in numerous applications, such as font design, biometric authentication, and robotic calli

Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.03071v1 Announce Type: new Abstract: Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting the right tool

Leveraging System-Level Observations to Inform Bayesian Learning of Model Parameters for Quantitative Verification

ApplicationsDGX agent

arXiv:2608.03489v1 Announce Type: cross Abstract: Combining Bayesian learning and quantitative verification is a powerful toolset for analysing key quantitative properties of software systems, like re

Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains

Model ReleasesDGX agent

arXiv:2608.03219v1 Announce Type: new Abstract: Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior. A model may

S^3-Diff: Structural Semantic Synergy Diffusion Model for High Fidelity Super Resolution of Pathological Images

ResearchDGX agent

arXiv:2608.03540v1 Announce Type: new Abstract: Digital pathology relies on high-resolution whole slide images for accurate diagnosis, yet limitations in imaging devices, storage, and transmission oft

Scalable Frequency- and Length-Aware Subdocument Deduplication for Large Language Model Pretraining

SafetyDGX agent

arXiv:2608.03089v1 Announce Type: new Abstract: Large-scale pretraining corpora contain substantial duplicate content. Although document-level deduplication is widely used, removing subdocument-level

Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling

Model ReleasesDGX agent

arXiv:2608.03842v1 Announce Type: new Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), 'which layer is responsible' has three natural operationalization

TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models

ResearchDGX agent

arXiv:2608.03057v1 Announce Type: new Abstract: Static quantization assigns one weight precision to every denoising step. To preserve quality, that precision must accommodate the most quantization-sen

4 Aug 2026

Agentic Graph Token Reasoning

AgentsDGX agent

arXiv:2608.00542v1 Announce Type: new Abstract: Graphs model relational data throughout science and industry, from citation networks to product co-purchase graphs. Because the nodes of many such graph

AROpt: An Optimization Method for Autoregressive Time Series Forecasting

ResearchDGX agent

arXiv:2602.02288v3 Announce Type: replace Abstract: Current time-series forecasting models are primarily based on transformer-style neural networks. These models achieve long-term forecasting mainly b

Beckmann Transport Models: From Autonomous Flows to One-Step Maps

AgentsDGX agent

arXiv:2608.01692v1 Announce Type: new Abstract: We propose an instantiation of flow matching that relies on a time-independent velocity field (an autonomous flow) to exactly map between two distributi

Climate-Dyna Deep Hedging for XVAs: Model-Based Reinforcement Learning, Residual Climate HVA, and Hedge-Instrument Discovery

ApplicationsDGX agent

arXiv:2608.01208v1 Announce Type: cross Abstract: For a trading desk, residual climate hedging valuation adjustment (HVA) is the climate cost left after its inherited hedge and any admissible overlay

Comparing and Modeling Argumentation in German Political Communication across Arenas

SafetyDGX agent

arXiv:2608.00288v1 Announce Type: new Abstract: Deliberation, involving the formulation and exchange of arguments, forms an integral part of political decision making in democracies. Argumentation pat

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

Model ReleasesDGX agent

arXiv:2608.01942v1 Announce Type: cross Abstract: Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

Model ReleasesDGX agent

First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that: https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_2

Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models

SafetyDGX agent

arXiv:2608.01263v1 Announce Type: new Abstract: On-policy distillation (OPD) samples trajectories from the current student policy and minimizes token-level divergence between student and teacher next-

DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable

Model ReleasesDGX agent

arXiv:2608.00548v1 Announce Type: new Abstract: Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their ras

DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation

SafetyDGX agent

arXiv:2608.01381v1 Announce Type: new Abstract: Mobile manipulation requires a robot to coordinate base and arm motion under continuously changing viewpoints and contact conditions, within an action s

Dynamic UAV-based search operations using probabilistic diffusion modeling of Man Overboard incident victims

ResearchDGX agent

arXiv:2608.02093v1 Announce Type: new Abstract: More than 70% of the people that fell overboard cruise ships in the period 2010-2019 lost their lives. This paper presents a strategy for reliably predi

Hugging Face CEO says China is winning the AI race and dominating on open models

Local AiDGX agent

This is something that was spoken here and there, and now it is like writing on the wall. The main additional point is that China has created an independent supply chain. Starting from raw materials a

Is LM Studio abandoning their core product?

Model ReleasesDGX agent

Some of you may be aware that a few weeks ago, LM Studio announced a new agent, Bionic. This is pretty much an agentic harness for both local models and paid cloud models. But most aren't aware that L

Parameter-Efficient CLIP Adaptation for 3D Understanding via Unified Tokenization

Model ReleasesDGX agent

arXiv:2505.18819v2 Announce Type: replace Abstract: Vision-language models, such as CLIP, encode rich semantic knowledge through large-scale image-text pretraining. Reusing these models for 3D underst

PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters' Lack of Knowledge

ResearchDGX agent

arXiv:2608.01598v1 Announce Type: new Abstract: Simulating human-like Theory of Mind (ToM) has been a longstanding problem in natural language processing (NLP). To address this, existing works introdu

PolymerGPT: Multi-property Optimization with a Decoder-Based GPT Model for Generative Polymer Design

ResearchDGX agent

arXiv:2608.01431v1 Announce Type: cross Abstract: Polymer property prediction and inverse generative design targeting desired properties are two crucial tasks in machine learning-assisted polymer desi

Provably Safe Generative Sampling with Constricting Barrier Functions

SafetyDGX agent

arXiv:2602.21429v3 Announce Type: replace Abstract: Flow-based generative models, such as diffusion models and flow matching models, have achieved remarkable success in learning complex data distribut

Quaternion Tensor Modeling for Joint Color-Polarization Demosaicking

ResearchDGX agent

arXiv:2608.02144v1 Announce Type: new Abstract: Division-of-focal-plane (DoFP) color polarization cameras enable snapshot acquisition of color polarization mosaic images, but the inherently sparse sam

Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks

Model ReleasesDGX agent

arXiv:2608.02442v1 Announce Type: cross Abstract: Scientific reasoning benchmarks typically evaluate large language models (LLMs) using final-answer accuracy. However, a correct answer does not necess

Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis

Local AiDGX agent

arXiv:2608.01973v1 Announce Type: cross Abstract: Existing indoor layout generators produce globally plausible layouts yet may retain local violations such as collisions, out-of-bounds placements, obs

Seeing the Unseen: Towards Training-Free Inspection for Wind Turbine Blades Using Knowledge-Augmented Vision Language Models

ResearchDGX agent

arXiv:2510.22868v2 Announce Type: replace Abstract: Wind turbine blades operate in harsh environments, making timely damage detection essential for preventing failures and optimizing maintenance. Dron

SPIRIT: Spatio-temporal Pairwise Relational Modeling of Instrument-Tissue Interactions for Surgical Action Triplet Recognition

SafetyDGX agent

arXiv:2608.02188v1 Announce Type: new Abstract: Fine-grained understanding of surgical activity is essential for context-aware assistance in the operating room, including safety monitoring, adverse ev

Surrogate Modeling for the Design of Optimal Lattice Structures using Tensor Completion

ResearchDGX agent

arXiv:2510.07474v2 Announce Type: replace Abstract: When designing new materials, it is often necessary to design a material with specific desired properties. Unfortunately, as new design variables ar

SVGEval: A Vision-Grounded Framework for Perceptual-Quality Benchmarking and Evaluation in Text-to-SVG Generation

Model ReleasesDGX agent

arXiv:2608.01977v1 Announce Type: new Abstract: Multimodal large models are increasingly used to generate scalable vector graphics (SVG), but reliable evaluation remains underexplored. Existing protoc

TabDPT-Turbo: Efficient In-Context Learning for Tabular Prediction

Model ReleasesDGX agent

arXiv:2608.01400v1 Announce Type: new Abstract: Tabular foundation models, driven by in-context learning, have rapidly grown in quality and popularity. However, recent approaches with either cell-base

Training nGPT

Model ReleasesDGX agent

arXiv:2608.01284v1 Announce Type: new Abstract: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the

Visualising Information Flow in Word Embeddings with Diffusion Tensor Imaging

ResearchDGX agent

arXiv:2601.05713v2 Announce Type: replace Abstract: Understanding how large language models (LLMs) represent natural language is a central challenge in natural language processing (NLP) research. Many

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

SafetyDGX agent

arXiv:2608.01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensi

3 Aug 2026

Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning

TutorialsDGX agent

arXiv:2607.19345v2 Announce Type: replace-cross Abstract: Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to lon

KAT Coder 2.5 dev: Do yourself a favor and try it!

Model ReleasesDGX agent

It is so good! I don't know why there aren't more people talking about it. Fewer tokens, faster and more accurate than Qwen 3.6 35b a3b. On my setup it's nearly as good as 27b, but 5x faster. And it c

OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems

Model ReleasesDGX agent

arXiv:2607.28629v1 Announce Type: new Abstract: The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the architectural und

ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding

Model ReleasesDGX agent

arXiv:2607.28751v1 Announce Type: new Abstract: Universal multimodal embedding (UME) maps heterogeneous multimodal inputs into a shared embedding space. Existing UME models either form embeddings thro

The Grokked Illusion: True Equilibrium Mitigates Catastrophic Forgetting

Model ReleasesDGX agent

arXiv:2607.29503v1 Announce Type: new Abstract: While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is

Token-Level Diagnosis of Sycophancy in LLMs with Attribution-Guided Steering

ResearchDGX agent

arXiv:2607.28906v1 Announce Type: new Abstract: Sycophancy refers to the tendency for large language models (LLMs) to match user beliefs at the cost of factual correctness, thereby undermining model r

Tokenizer-Agnostic Engram Module

Model ReleasesDGX agent

arXiv:2607.29065v1 Announce Type: new Abstract: Deepseek's Engram, a conditional memory module, was introduced to trade-off storage versus reasoning in large language models. However, the module relie

2 Aug 2026

DeepSeek-V4-Flash 284B on 5.3GB of memory

Model ReleasesDGX agent

Following up on my Qwen 3.6 port, I wanted to keep adding models and ended up fixing a bunch of things along the way, so it's its own engine now: Mference. Same core idea from TurboFieldfare, MoE mode

How well do multiple GPUs scale for LLM inference? (Trying to understand the basics)

Model ReleasesDGX agent

Hi everyone, I’m fairly new to the multi-GPU side of local LLMs and I’m trying to understand how inference actually scales across multiple GPUs. Suppose I have a model running on a single GPU and then

Real-world reality check on Qwen for autonomous coding agents

Model ReleasesDGX agent

TLDR below 👇🏼 I’ve seen a lot of hype around Qwen 3.6 35B and 3.5 120B lately, especially regarding coding and tool-use capabilities. On this subreddit it is the defacto recommended model for everyone

1 Aug 2026

There's no 'one weird trick” for prompting Krea 2 art styles—just many guidelines [WF included]

TutorialsDGX agent

TLDR: There is no one prompting trick that will result in Krea 2 Turbo giving you exactly the style you want and across the whole image. Instead, if you are trying to achieve styles without the use of

31 Jul 2026

AI-native software development requires a new engineering model

IndustryDGX agent

Artificial intelligence has quickly become a standard part of modern software development. Coding assistants, code completion tools and AI-powered integrated development environments are now widely av

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

Model ReleasesDGX agent

arXiv:2607.28146v1 Announce Type: new Abstract: As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabil

Co-Designing AI Model Attention for Fast, Interactive Long-Context Inference

HardwareDGX agent

The article shows that dense‑attention performance in long‑context inference is governed by group size (query heads per KV head), head dimension, and sequence length, with prefill being compute‑bound

IGME: Efficient Chained Method Ensemble for Transferable Semantic Segmentation Attacks

ResearchDGX agent

arXiv:2607.27465v1 Announce Type: new Abstract: Semantic segmentation models are vulnerable to transferable adversarial perturbations, yet evaluating transfer attacks on dense prediction models can be

Latent-Kernel Discrete Flow Maps for Few-Step Generation

ResearchDGX agent

arXiv:2607.27529v1 Announce Type: new Abstract: Discrete diffusion and flow-matching models denoise a sequence over many steps, but to keep each step cheap, they factorize the transition across positi

MORFES: A Benchmark for Productive Inflectional Competence in Modern Greek

Model ReleasesDGX agent

arXiv:2607.28274v1 Announce Type: new Abstract: Modern Greek is a richly inflected language, yet the language models built for it are evaluated mainly on factual knowledge, and no benchmark is dedicat

Scalable Drift Monitoring in Medical Imaging AI

ApplicationsDGX agent

arXiv:2410.13174v3 Announce Type: replace-cross Abstract: The integration of artificial intelligence (AI) into medical imaging has advanced clinical diagnostics but poses challenges in managing model

Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

Model ReleasesDGX agent

arXiv:2607.27232v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory

HardwareDGX agent

arXiv:2607.28263v1 Announce Type: new Abstract: Transformer depth is not used uniformly: lower and middle layers build semantic representations, while upper layers increasingly specialize them for pre

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

ResearchDGX agent

arXiv:2607.26742v1 Announce Type: cross Abstract: Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual informatio

30 Jul 2026

Benchmarked: MindControl for Llama.cpp

Model ReleasesDGX agent

I recently shared the original MindControl PoC (and on github) - sampler-level guided reasoning budgets for llama.cpp, nudging the model with self-aware statements about its own thinking budget instea

CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

SafetyDGX agent

arXiv:2607.26789v1 Announce Type: new Abstract: Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions withou

← Previous
1…245246247248249…1018
Next →