AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,292 results
5 Aug 2026

Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition

Model ReleasesDGX agent

arXiv:2608.03892v1 Announce Type: new Abstract: We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, r

Introducing Muse Code and Muse Spark 1.2

Model ReleasesDGX agent

Introducing Muse Code and Muse Spark 1.2 Yet more evidence that the most important characteristic of any model these days is long-sequence agentic tool calling. Meta shipped their own coding agent as

Inverted Detection and Control in Steering Vectors

Model ReleasesDGX agent

arXiv:2608.02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.g., truthfulness) in large language model outputs. A key assumption un


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning

Model ReleasesDGX agent

arXiv:2507.14171v3 Announce Type: replace-cross Abstract: Importance-based structured pruning overwhelmingly relies on filter magnitude. This proxy is fundamentally flawed: due to scale invariance, fu

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion

Model ReleasesDGX agent

arXiv:2608.03974v1 Announce Type: new Abstract: Real-time video editing requires low-latency causal generation with bounded computational resources while preserving source fidelity and long-term tempo

JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation

Model ReleasesDGX agent

arXiv:2608.02620v1 Announce Type: new Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their o

KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation

Model ReleasesDGX agent

arXiv:2608.03782v1 Announce Type: new Abstract: Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks mainly focus o

Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks

Model ReleasesDGX agent

arXiv:2608.02621v1 Announce Type: cross Abstract: Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy fo

LAEF: A Lead-Agnostic ECG Foundation Model Towards Point-of-Care Diagnostics

Model ReleasesDGX agent

arXiv:2608.03690v1 Announce Type: new Abstract: Point-of-care cardiac devices such as smartwatches and handheld ECG recorders typically capture 1--2 leads, yet existing ECG foundation models are archi

Latent Reward Registers for Diffusion Preference Alignment

Model ReleasesDGX agent

arXiv:2608.03929v1 Announce Type: cross Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a sev

LDU-Bench: Multimodal LLM Evaluation for Lithography Defect Understanding under Layout-Varying Circuit Backgrounds

Model ReleasesDGX agent

arXiv:2608.03078v1 Announce Type: new Abstract: Multimodal large language models have demonstrated strong defect recognition capability in industrial anomaly detection. However, in lithography review,

LeanMem: Simple and Efficient Long-Term Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2608.03463v1 Announce Type: new Abstract: Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history. However, existing memory systems typic

Let's create with Qwen-Image-3.0-Pro on fal! 🎨

Model ReleasesDGX agent

Let's create with Qwen-Image-3.0-Pro on fal! 🎨 Qwen Image 3.0 Pro is now available on fal - Renders complex, dense typography - Preserves key details like facial features and identity while applying c

LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU

Model ReleasesDGX agent

As you all know the model is 2.69B parameters with a 128K context window and purpose-built for multi-step agent workflows. What you are seeing is the Q4_K_M GGUF running on my own inference engine bui

Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2608.03148v1 Announce Type: cross Abstract: RAG improves the factual grounding of LLM by incorporating external knowledge, but deploying RAG on mobile and edge devices remains challenging becaus

LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment

Model ReleasesDGX agent

arXiv:2608.03020v1 Announce Type: new Abstract: Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation through the frozen

Localize, Don't Beautify: Client-Side Control of Image-Editing APIs for Cosmetic Surgery Previews

Model ReleasesDGX agent

arXiv:2608.02841v1 Announce Type: new Abstract: Ask a commercial image editor to preview a cosmetic procedure and it will often change more of the face than the request names: a nose edit can also smo

LocAnyMed: Vision-Language Grounding for Multimodal Medical Images

Model ReleasesDGX agent

arXiv:2608.03322v1 Announce Type: new Abstract: Medical visual grounding connects free-form clinical queries to spatial evidence in medical images and is an important component of interpretable medica

Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation

Model ReleasesDGX agent

arXiv:2608.03577v1 Announce Type: new Abstract: Automation of Translation Quality Estimation (QE) has emerged as a widely discussed approach to managing translation quality at scale, and a growing num

LoopMTP: A looped transformer guided by latent multi-token prediction

Model ReleasesDGX agent

arXiv:2608.03624v1 Announce Type: new Abstract: Looped transformers have emerged as a parameter-efficient alternative to scaling depth for strong reasoning. By reusing one stack of layers across T ite

Low-Dimensional High-Leverage Subspace Optimization: Beyond Full-Parameter Coupled Training for Neural Network Quantization

Model ReleasesDGX agent

arXiv:2608.03919v1 Announce Type: new Abstract: Low-bit quantization suffers severe accuracy degradation on compact networks, rooted in the dominant full-parameter coupled training paradigm that ignor

M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2608.03803v1 Announce Type: new Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a languag

Maglev: Sliding Recurrent Memory

Model ReleasesDGX agent

arXiv:2608.02870v1 Announce Type: new Abstract: We introduce ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizabl

Material-Segmented Per-Pixel Emissivity Correction for Thermographic Anomaly Detection in Cultural Heritage Digital Twins

Model ReleasesDGX agent

arXiv:2608.02964v1 Announce Type: new Abstract: Quantitative longwave thermography of heritage surfaces is limited by the global-constant emissivity assumption in inverse-Planck temperature retrieval;

MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows

Model ReleasesDGX agent

arXiv:2608.02642v1 Announce Type: cross Abstract: Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out as a partic

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

Model ReleasesDGX agent

arXiv:2608.00794v2 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims.

MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

Model ReleasesDGX agent

arXiv:2608.02613v1 Announce Type: cross Abstract: Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory bench

MIDI-LLM: Improving Text-to-MIDI Music Generation via Adapting Large Language Models

Model ReleasesDGX agent

arXiv:2511.03942v2 Announce Type: replace-cross Abstract: We present MIDI-LLM, a recipe that improves multitrack text-to-MIDI generation via adapting Large Language Models (LLMs). MIDI-LLM expands an

MinerU.Chem: A High-Precision System for Optical Chemical Structure and Reaction Recognition

Model ReleasesDGX agent

arXiv:2608.03525v1 Announce Type: new Abstract: In organic chemistry papers and patents, molecular structures, reaction schemes, and experimental conditions are often presented as molecular structure

Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue

Model ReleasesDGX agent

arXiv:2608.03142v1 Announce Type: cross Abstract: We study contextual dynamic pricing with arbitrary covariate sequences and bounded, possibly nonbinary purchase quantities. Demand follows a semiparam

Mistral introduces Shieldstral to provide lightweight policy-aware moderation for AI models

Model ReleasesDGX agent

French artificial intelligence startup Mistral AI SAS today introduced a lightweight multimodal safety artificial intelligence open-weight model that can classify outputs for AI models that outperform

MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc

Model ReleasesDGX agent

arXiv:2608.03397v1 Announce Type: new Abstract: MMLongBench-Doc is a long-document QA benchmark of 1,082 questions over 135 PDFs. Two properties of it push measured scores away from the quantity they

Modeling Long-Term Memory and Temporal Attention Shifts for Video Salient Object Ranking with a New Benchmark

Model ReleasesDGX agent

arXiv:2203.17257v2 Announce Type: replace Abstract: Salient Object Ranking (SOR) aims to estimate the relative saliency order among multiple salient objects. While SOR has been extensively studied in

MoE CPU-offload benchmark on Deepseek V4/Gemma4/Qwen/GPT-OSS — TensorSharp vs llama.cpp

Model ReleasesDGX agent

TensorSharp's MoE CPU-offload feature has been merged into main. Here is the parameters description of this feature: Mixture-of-Experts CPU offload: --n-cpu-moe <N> | -ncmoe <N> Keep the routed MoE ex

MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation

Model ReleasesDGX agent

arXiv:2608.03275v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capa

Morphology-Aware Implicit Super-Resolution Network for Pathological Images

Model ReleasesDGX agent

arXiv:2608.03664v1 Announce Type: new Abstract: Accurate diagnosis in Digital Pathology (DP) relies on high-resolution whole-slide images, yet clinical deployment is often limited by hardware costs. S

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification

Model ReleasesDGX agent

arXiv:2608.03474v1 Announce Type: new Abstract: Recent advances in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in web UI generation. However, existing benchmarks pre

MultiCompose: Multi-Concept Personalized Composition with Per-Subject Attribute Binding

Model ReleasesDGX agent

arXiv:2608.03708v1 Announce Type: new Abstract: Text-to-image diffusion models enable personalization of specific visual concepts from a small number of reference images. However, generating a single

MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

Model ReleasesDGX agent

arXiv:2608.03882v1 Announce Type: cross Abstract: Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logis

MuRA: Multi-Rank Adaptation for Efficient and Effective Test-Time Vision-Language Generalization

Model ReleasesDGX agent

arXiv:2608.03885v1 Announce Type: new Abstract: Vision-language models exhibit remarkable zero-shot capabilities but suffer significant performance degradation under distribution shifts. While test-ti

NearID: Identity Representation Learning via Near-identity Distractors

Model ReleasesDGX agent

arXiv:2604.01973v2 Announce Type: replace Abstract: When evaluating identity-focused tasks such as personalized generation and image editing, existing vision encoders entangle object identity with bac

Neural Networks with Local Converging Inputs for Efficient Options Pricing Models

Model ReleasesDGX agent

arXiv:2608.02778v1 Announce Type: new Abstract: We present a novel application of Neural Networks with Local Converging Inputs (NNLCI) to improve the efficiency of existing numerical methods for prici

NOMADD: Numerical Optimization of Models Adapting to Data Drift

Model ReleasesDGX agent

arXiv:2608.02845v1 Announce Type: new Abstract: Tabular model performance degrades when feature distributions change over time or the relationship between features and outcome variables change over ti

Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes

Model ReleasesDGX agent

arXiv:2608.03839v1 Announce Type: new Abstract: Training-free speculative decoding drafts by matching an exact suffix of the context against a pool of earlier context. That lookup misses correct draft

OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet

Model ReleasesDGX agent

arXiv:2608.03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due

Omega-S: A Functional Resilience Index for LLM Fine-Tuning

Model ReleasesDGX agent

arXiv:2608.03887v1 Announce Type: new Abstract: Fine-tuning a large language model on new data degrades what it previously learned. We present Omega-S, a drop-in penalty computed from the weight matri

On the Implicit Flatness Bias of Sharpness-Aware Minimization: A Linear Stability Analysis with Quantitative Hyperparameter Bounds

Model ReleasesDGX agent

arXiv:2608.03197v1 Announce Type: new Abstract: Sharpness-Aware Minimization (SAM) improves generalization by seeking parameters whose loss is robust to local adversarial perturbations, but the quanti

On the missing benchmarks layer and a potential solution

Model ReleasesDGX agent

arXiv:2608.02996v1 Announce Type: new Abstract: Latin America is missing a foundational layer for native AI development: the benchmark layer. The benchmark layer does two things no other layer can - i

On the missing data layer and a potential solution

Model ReleasesDGX agent

arXiv:2608.02949v1 Announce Type: new Abstract: Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the dataset layer.

OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning

Model ReleasesDGX agent

arXiv:2608.02615v1 Announce Type: cross Abstract: Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata. However,

One Knob to Rule Them All: A Unified Optimal Transport View of Cold-Start Active Learning

Model ReleasesDGX agent

arXiv:2608.03249v1 Announce Type: new Abstract: Cold-Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance. Existing meth

One-Point Contraction: Erasing Representational Separability toward Irreversible Deep Forgetting

Model ReleasesDGX agent

arXiv:2507.07754v3 Announce Type: replace-cross Abstract: Machine unlearning is usually evaluated by what the classifier outputs: forget-set accuracy, confidence, membership-inference scores. We show

One-shotting a Raccoon Heist game using Claude Fable 5

Model ReleasesDGX agent

Back in 2024 I tweeted screenshots of a game concept generated by GPT-3 and some concept 'art' created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5

Operationally Feasible Synthetic Power-Grid Scenarios via Learning the AC-Operable Joint Distribution

Model ReleasesDGX agent

arXiv:2608.03878v1 Announce Type: new Abstract: Synthetic power-grid scenarios are essential for planning, resilience assessment, contingency analysis, and data-driven power-system applications. Recen

PACE: Adaptive Budget Allocation for Time-Efficient Embodied Planning

Model ReleasesDGX agent

arXiv:2608.03034v1 Announce Type: cross Abstract: Reasoning-enhanced large language models have achieved remarkable improvements in planning tasks, yet their deployment in embodied systems remains imp

Particle-based Generalised Stochastic Optimisation

Model ReleasesDGX agent

arXiv:2608.02844v1 Announce Type: cross Abstract: We develop a class of diffusion-based stochastic particle optimisation methods for loss functions with intractable gradients. Specifically, we conside

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

Model ReleasesDGX agent

arXiv:2608.04010v1 Announce Type: cross Abstract: Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation,

PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents

Model ReleasesDGX agent

arXiv:2608.04003v1 Announce Type: new Abstract: Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for s

Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation

Model ReleasesDGX agent

arXiv:2608.03691v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are increasingly used to translate webpage screenshots into front-end code, but repeated UI patterns may sway

PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory

Model ReleasesDGX agent

arXiv:2608.03048v1 Announce Type: cross Abstract: Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent challenges: se

← Previous
1…3031323334…372
Next →