AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

86,457Total entries
1Added by human
86,456Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,039 results
28 Jul 2026

Scale Weight Decay and Train Better

Model ReleasesDGX agent

arXiv:2607.23777v1 Announce Type: cross Abstract: The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data. This is typically done with a constant dec

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling

Model ReleasesDGX agent

arXiv:2510.14717v2 Announce Type: replace-cross Abstract: Increasing the batch size during training -- a ''batch ramp'' -- is a promising strategy to accelerate large language model pretraining. While

Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning

Model ReleasesDGX agent

arXiv:2607.22732v1 Announce Type: new Abstract: LLM-based game agents often perform poorly on more complex tasks. This work examines whether these failures are linked to limited spatial reasoning and

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting

Model ReleasesDGX agent

arXiv:2607.24191v1 Announce Type: cross Abstract: Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit three key l

SWE-rebench Multilingual Update (Go, Java, Python, Rust, TS). Evaluated: GLM-5.2, DeepSeek-V4 Pro, Qwen3.6-27B and others

Model ReleasesDGX agent

Hi everyone! We’ve just released a major update to the leaderboard! We are expanding beyond Python with a new multilingual slice featuring real-world software engineering tasks across 5 languages. Ope

Tailored untruths: How personalisation challenges LLM safeguards

SafetyDGX agent

arXiv:2510.12993v3 Announce Type: replace Abstract: Large Language Models (LLMs) can generate highly persuasive disinformation, yet little is known about how effectively they personalise it across lan

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

SafetyDGX agent

arXiv:2607.24720v1 Announce Type: cross Abstract: Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are tra

Verbalized Particle Posterior: Bayesian Inference over Natural Language Hypotheses

Model ReleasesDGX agent

arXiv:2607.22961v1 Announce Type: cross Abstract: Verbalized Machine Learning (VML) parameterizes a model as a natural-language prompt that an LLM evaluates as f(x; theta). The framework is interpreta

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

Model ReleasesDGX agent

arXiv:2607.24392v1 Announce Type: cross Abstract: Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. W

WISERouter: LLM Routing with Workload Budget Constraint

ResearchDGX agent

arXiv:2607.23765v1 Announce Type: cross Abstract: Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is prohibitive a

27 Jul 2026

b10153

Model ReleasesDGX agent

model: Add support for Nanbeige4.2 (#25994) support nanbeige4.2 model fix fix flake8 Lint check fix loop bound check and drop redundant head_dim Co-authored-by: root lizongqiang@kanzhun.com Website: h

DM3D: Dynamic Mamba via Offset-Guided Feature Resampling for Point Cloud Understanding

Model ReleasesDGX agent

arXiv:2512.03424v4 Announce Type: replace Abstract: State Space Models (SSMs) model long token sequences of point cloud with linear complexity, but require an unordered point cloud to be serialized. E

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision

Model ReleasesDGX agent

arXiv:2603.07025v2 Announce Type: replace Abstract: Speech Large Language Models (LLMs) that understand and follow instructions in many languages are useful for real-world interaction, but are difficu

LMEB: Long-horizon Memory Embedding Benchmark

Model ReleasesDGX agent

arXiv:2603.12572v5 Announce Type: replace Abstract: Memory embeddings are crucial for memory-augmented systems, such as OpenClaw, but their evaluation is underexplored in current text embedding benchm

Offline Vision-Language Navigation with Geometric Goal Localization for Outdoor Environments

Model ReleasesDGX agent

arXiv:2607.22226v1 Announce Type: new Abstract: Foundation-model-based vision-language navigation (VLN) has advanced autonomous robot navigation by enabling robots to interpret natural-language instru

Optimization of time-consuming experimental conditions using pseudo-experimental data guided by adaptive polynomial regression

Model ReleasesDGX agent

arXiv:2607.22238v1 Announce Type: new Abstract: Bayesian optimization (BO) is an optimization method that sequentially proposes the next candidate explainable variables for optimizing target variables

RadSight: Towards Perceptually Reliable Multimodal Radiology Image Understanding

Model ReleasesDGX agent

arXiv:2607.22293v1 Announce Type: new Abstract: Medical multimodal large language models (MLLMs) are increasingly expected to perform complex image understanding tasks, yet their reliability is often

Scaling Laws for Classical Machine Learning on Tabular Data: A Benchmark Study

Model ReleasesDGX agent

arXiv:2607.21866v1 Announce Type: new Abstract: Prior classical-ML learning-curve work fits power laws to tree, linear, and kernel models on tabular data, but at small scale: typically one curve, one

Twins: Learn to Predict Unified Representations with Focal Loss

SafetyDGX agent

arXiv:2607.22531v1 Announce Type: new Abstract: Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the

Unbiased Open World Regularization for Fair Self-Supervised Learning

Model ReleasesDGX agent

arXiv:2607.22149v1 Announce Type: new Abstract: Despite recent advances, self-supervised learning (SSL) models and Joint-Embedding Predictive Architectures (JEPAs) remain susceptible to learning spuri

26 Jul 2026

ai-sage/GigaChat3.1-Audio-10B-A1.8B · Hugging Face

Local AiDGX agent

GigaChat Audio 10B is an audio-native LLM built on top of the GigaChat 3.1 Lightning text model. A Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Expe

This part is spot on: > 'Overall, we found that we were over-constraining Claude Code...while these constraints were once needed to avoid wo…

Model ReleasesDGX agent

This part is spot on: > 'Overall, we found that we were over-constraining Claude Code...while these constraints were once needed to avoid worst case scenarios, we have since found we can delete many o

25 Jul 2026

Deepseek V4 flash - Hy3 or is Qwen3.6 27B still the most solid for agentic/coding?

Model ReleasesDGX agent

I understand that the laguna model is either still buggy or potentially benchmaxxed. So I’d like to know for people who really tested, are DS flash or Hy3 really better in your usecase? submitted by /

Kimi Linear 48B A3B?

Model ReleasesDGX agent

Just noticed this exists, 1M context MOE with 48B par seems just like what Ive been looking for - it runs pretty damn fast too compared to Qwen 3.6 35B. after some testing it seems capable of producin

24 Jul 2026

AI-Driven Surrogate Models for Predicting Electrode-Scale Discharge Behavior in Lithium-Ion Batteries

ResearchDGX agent

arXiv:2607.20577v1 Announce Type: new Abstract: Physics-based simulations are essential for understanding the electrode-scale discharge behavior of lithium-ion batteries (LIBs) but suffer from prohibi

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

Model ReleasesDGX agent

arXiv:2607.20596v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE fami

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

Model ReleasesDGX agent

arXiv:2607.20465v1 Announce Type: cross Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how

Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks

Model ReleasesDGX agent

arXiv:2607.20820v1 Announce Type: new Abstract: Body-based emotion recognition is important for real-time affective systems, but graph-based skeleton models can be computationally expensive. This pape

Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time

Model ReleasesDGX agent

arXiv:2603.20509v2 Announce Type: replace Abstract: Camera traps are vital for large-scale biodiversity monitoring, yet accurate automated analysis remains challenging due to diverse deployment enviro

Memoir: Should a Model Write to Its Memory While It Thinks?

ResearchDGX agent

arXiv:2607.20792v1 Announce Type: new Abstract: Memoir combines per-sample fast memory, shared slow parameters, variable-depth latent recurrence, and a future-latent energy objective. We test its risk

NVIDIA, Palantir, Replit, Microsoft, Crowdstrike, Dell and others send a strong message to congress to keep open access to open weights mode…

SafetyDGX agent

NVIDIA, Palantir, Replit, Microsoft, Crowdstrike, Dell and others send a strong message to congress to keep open access to open weights models. 'Our AI leadership will be judged not by one frontier AI

Post-Hoc Reasoning in Chain of Thought: Decoding and Steering Pre-Committed Answers

ResearchDGX agent

arXiv:2603.01437v2 Announce Type: replace Abstract: As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising t

Scaling Closed-Loop Feature Channel Configuration with LLMs

Model ReleasesDGX agent

arXiv:2607.20516v1 Announce Type: cross Abstract: Promising initial results in closed-loop large-language-model-based channel-configuration search demonstrated that neural-network widths can be optimi

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Model ReleasesDGX agent

arXiv:2607.21072v1 Announce Type: new Abstract: Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks a

The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs

SafetyDGX agent

arXiv:2607.20449v1 Announce Type: cross Abstract: LLMs are trained predominantly on human-authored text, yet the structural and narrative conventions embedded in that text are rarely examined as a sou

23 Jul 2026

Associative Emotional Learning in Convolutional Neural Networks

ResearchDGX agent

arXiv:2607.19327v2 Announce Type: replace Abstract: Associative emotional learning enables organisms to adaptively link pleasant or unpleasant outcomes to the presence of predictive stimuli. Whereas c

Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation

Local AiDGX agent

arXiv:2602.19863v3 Announce Type: replace Abstract: Foundation models are transforming Earth Observation (EO), yet the diversity of EO sensors and modalities makes a single universal model unrealistic

Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting

Model ReleasesDGX agent

arXiv:2607.20086v1 Announce Type: cross Abstract: State-space sequence models are attractive for streaming speech because they maintain compact recurrent state, but scan-style training kernels can hav

Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2607.19539v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block

GH-ESD: Grounded Hypothesis-Driven Error Slice Discovery for Instance-Level Vision Tasks

Model ReleasesDGX agent

arXiv:2512.24592v3 Announce Type: replace Abstract: Systematic failures of vision models on semantically coherent subsets, known as error slices, reveal limitations in robustness and evaluation. Exist

Group-of-Latents: Perceptual Video Compression at Extreme Bitrates via Masked Latent Generative Modeling

ResearchDGX agent

arXiv:2607.19437v1 Announce Type: cross Abstract: Most existing video compression algorithms follow a paradigm of transformation and quantization, optimizing the trade-off between distortion and bitra

How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF

SafetyDGX agent

arXiv:2607.19712v1 Announce Type: new Abstract: In RLHF pipelines, reward scoring blocks policy updates. Slow scoring bottlenecks the entire loop, since no update runs until every rollout gets a score

MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation

Model ReleasesDGX agent

arXiv:2607.14595v2 Announce Type: replace Abstract: Large-scale video diffusion models deliver strong generation performance, but full fine-tuning for downstream tasks incurs prohibitive computational

OffNadirLoc: Benchmark and Framework for Challenging UAV-to-Satellite Geo-Localization under Large Off-Nadir Views

Model ReleasesDGX agent

arXiv:2607.19951v1 Announce Type: new Abstract: Cross-view geo-localization between UAV and satellite imagery remains a fundamental yet highly challenging task, especially under large off-nadir views

Solar Open 2 Technical Report

Model ReleasesDGX agent

arXiv:2607.20062v1 Announce Type: new Abstract: We present Solar Open 2, a 250B-A15B Mixture-of-Experts language model built for long-horizon agentic tasks, scaled up from Solar Open 1 (Solar Open 100

Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time Series Forecasting

SafetyDGX agent

arXiv:2607.19404v1 Announce Type: cross Abstract: Multivariate time series encode structural patterns that unfold across multiple temporal scales, yet most forecasting backbones treat learned represen

The first known runaway AI agent - or a very bad marketing stunt?

Model ReleasesDGX agent

The first known runaway AI agent - or a very bad marketing stunt? Martin Alderson's commentary on the OpenAI accidental cyberattack against Hugging Face includes a couple of details I hadn't considere

22 Jul 2026

Are AI labs pelicanmaxxing?

Model ReleasesDGX agent

Are AI labs pelicanmaxxing? Excellent piece of work by Dylan Castillo, who took a deep-dive into the frequently pondered question of whether the AI labs have been deliberately training models to draw

Instead of panicking about the Hugging Face attack, people need to start questioning OpenAI's insecure sandboxes.

Model ReleasesDGX agent

One thing I noticed in American politics, whenever the government wants to push unpopular actions or laws, they often introduce fear to convince the public to support them. This is actually how i view

MindControl - llama.cpp fork to guide the reasoning process via injection during sampling

Model ReleasesDGX agent

The primary driver of this project is that I'd become frustrated with the reasoning behavior of smaller local models such as Qwen3.6-27B (i believe particularly at lower temperatures, and where system

21 Jul 2026

Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Model ReleasesDGX agent

Google released three new Gemini AI models on July 21, 2026: Gemini 3.6 Flash, Gemini 3.5 Flash‑Lite, and Gemini 3.5 Flash Cyber. These models are designed to deliver higher token‑efficiency, lower la

20 Jul 2026

Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are…

Model ReleasesDGX agent

Exciting update: Kimi K3 has landed at #4 on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1 open

16 Jul 2026

Anatomically Faithful but Temporally Blind: Auditing Attribution for Left-Ventricular Ejection-Fraction Estimation from Echocardiography

Local AiDGX agent

arXiv:2607.13738v1 Announce Type: cross Abstract: Background and Objective: Deep video models estimate left-ventricular ejection fraction (EF) from echocardiography with near-expert accuracy, and post

Efficient Text-to-Audio Generation via Pruning

Model ReleasesDGX agent

arXiv:2607.13330v1 Announce Type: cross Abstract: Diffusion-based text-to-audio generative models such as AudioLDM achieve high perceptual quality and strong semantic consistency; however, their pract

Evaluating Frontier AI Agents as Autonomous Clinical Security Auditors

Model ReleasesDGX agent

arXiv:2607.13411v1 Announce Type: cross Abstract: Clinical AI models can expose patients to harm when adversarial vulnerabilities go undetected, yet formal security auditing requires statistical exper

Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate

Model ReleasesDGX agent

arXiv:2510.10002v3 Announce Type: replace Abstract: As large language models (LLMs) are increasingly deployed in sensitive everyday contexts -- offering personal advice, mental health support, and mor

Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots

ResearchDGX agent

arXiv:2607.13522v1 Announce Type: cross Abstract: A robot must understand the state of its own body, but a camera sees only part of it. Force and contact leave almost no trace in a single frame, and r

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B on 2x3090s

Model ReleasesDGX agent

I managed to get this model working on 2x 3090s with full 262k ctx and N=4, if anyone is interested to try it, thanks to this quant: https://huggingface.co/danielrmay/NVIDIA-Nemotron-Labs-3-Puzzle-75B

OvisOCR2 Technical Report

Model ReleasesDGX agent

arXiv:2607.13639v1 Announce Type: cross Abstract: We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generates a Markdo

Policy of Thoughts: Scaling Test-Time Training for LLM Reasoning via Online Policy Evolution

Model ReleasesDGX agent

arXiv:2601.20379v2 Announce Type: replace Abstract: Large language models (LLMs) struggle with complex, long-horizon reasoning due to instability caused by their frozen policy assumption. Current test

← Previous
1…273274275276277…1034
Next →