AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
4 Aug 2026

Decoupling semantics from vision: A framework for faithful visual-text compression evaluation

Model ReleasesDGX agent

arXiv:2608.01848v1 Announce Type: new Abstract: Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks

DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

Model ReleasesDGX agent

Just wanted to share my agentic coding benchmark run of DSv4F 0731 at both High and Low reasoning efforts (not Max)... I ran a 109-question subset of Aider Polyglot (the JS/C++/Python languages), base

Defending Membership Inference Attacks via Privacy-aware Sparsity Tuning

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2410.06814v2 Announce Type: replace Abstract: Over-parameterized models are typically vulnerable to membership inference attacks, which aim to determine whether a specific sample is included in

DeltaFlow: Noise-Adaptive Bidirectional Gated Delta Networks for Embedded Language Flows

Model ReleasesDGX agent

arXiv:2608.01240v1 Announce Type: new Abstract: Embedded Language Flows (ELF) rely primarily on full non-causal attention for iterative denoising, repeatedly incurring quadratic sequence-mixing cost a

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs

Model ReleasesDGX agent

arXiv:2608.01979v1 Announce Type: cross Abstract: Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In

Feed-Forward Steering in Transformer Residual Dynamics

Model ReleasesDGX agent

arXiv:2608.02071v1 Announce Type: new Abstract: Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere. We extend this framework by incorporating

GPT-OSS has turned one year old today!

Model ReleasesDGX agent

It is one of the best local models ever released, in both 20B and 120B versions. I always come back to it, especially the 120B version. Its only competition is, in my opinion, Qwen 3.5 122B, but that

Grounding Agentic VLMs with Dedicated Segmentation for Fine-Grained Vehicle Damage Assessment

Model ReleasesDGX agent

arXiv:2608.02470v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed as reasoning agents in real-world visual assessment pipelines, yet their spatial grounding remai

HarMoE: Multi-Source Chest Radiograph Pretraining with Dataset-Disentangled Experts

Model ReleasesDGX agent

arXiv:2608.02252v1 Announce Type: new Abstract: Recent vision-language models for chest X-ray understanding are largely built on image-report alignment and therefore rely heavily on MIMIC-CXR as the d

Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races

Model ReleasesDGX agent

arXiv:2608.01193v1 Announce Type: cross Abstract: An AI development race creates a multi-agent safety dilemma. Each company can develop slowly and safely, or move faster while taking a risk that may r

InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos

Model ReleasesDGX agent

arXiv:2608.01157v1 Announce Type: new Abstract: Large language models have made text the default medium for human--AI interaction, buttext alone cannot express the full range of responses required by

KING: Embodiment-Aware Kinematic Graph Neural Network for Unified Motion Representation of Legged and Wheeled Robots

ResearchDGX agent

arXiv:2608.01015v1 Announce Type: new Abstract: Kinematic models provide reliable motion constraints for odometry estimation in featureless environments, where exteroceptive sensing degrades and IMU i

LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering

Model ReleasesDGX agent

arXiv:2604.03532v2 Announce Type: replace Abstract: Large language models (LLMs) show strong multilingual capabilities, yet reliably controlling the language of their outputs remains difficult. Repres

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

Local AiDGX agent

arXiv:2608.01395v1 Announce Type: new Abstract: We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU la

llm-anthropic 0.26

Model ReleasesDGX agent

Release: llm-anthropic 0.26 Includes new features enabled by LLM 0.32: New models: claude-fable-5, claude-sonnet-5, and claude-opus-5. #75, #76 Added server-side tools for WebSearch, WebFetch, CodeExe

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

Model ReleasesDGX agent

arXiv:2608.01964v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interde

Machine-Precision Prediction of Low-Dimensional Chaotic Systems from Noise-Free Data

Model ReleasesDGX agent

arXiv:2507.09652v2 Announce Type: replace-cross Abstract: Low-dimensional chaotic systems such as the Lorenz-63 model are commonly used to benchmark system-agnostic methods for learning dynamics from

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

Model ReleasesDGX agent

arXiv:2608.02595v1 Announce Type: new Abstract: Language models are playing an increasingly important role in laboratory science, performing tasks such as experiment planning, execution, and post-hoc

OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

Model ReleasesDGX agent

arXiv:2608.00677v1 Announce Type: new Abstract: AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model i

Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel

Model ReleasesDGX agent

arXiv:2608.00979v1 Announce Type: cross Abstract: Large language models are increasingly used as synthetic research participants and are often validated by whether their marginal responses resemble hu

PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs

ResearchDGX agent

arXiv:2608.02150v1 Announce Type: new Abstract: Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an understanding of

RAP: KV-Cache Compression via RoPE-Aligned Pruning

Model ReleasesDGX agent

arXiv:2602.02599v4 Announce Type: replace Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the memory and compute of the key-value (KV) cache. Structured pruning is

Remember OpenAI tripled its score on ARC-AGI with a harness fix? We fixed harness for Kimi K3 on CyberGym E2E: 2x better at vulnerability de…

Model ReleasesDGX agent

Remember OpenAI tripled its score on ARC-AGI with a harness fix? We fixed harness for Kimi K3 on CyberGym E2E: 2x better at vulnerability detection, 3x at patching K3 is SOTA cyber-defense model you c

S^4R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching

Model ReleasesDGX agent

arXiv:2608.00528v1 Announce Type: new Abstract: The growth of context window lengths in Large Language Models (LLMs) significantly enhances their long-context capabilities but incurs prohibitive memor

SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling

Model ReleasesDGX agent

arXiv:2608.00991v1 Announce Type: cross Abstract: This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form vari

SLMs as Multi-Agent Routers: A Progressive SFT and Reinforcement Learning Approach

Model ReleasesDGX agent

arXiv:2608.00030v1 Announce Type: new Abstract: Specialised retrieval agents typically surface higher quality results than general-purpose search, but selecting the optimal agent for a given query rem

Test-Time Curriculum for Open-Set AIGC Detection

Model ReleasesDGX agent

arXiv:2608.00559v1 Announce Type: new Abstract: AI-generated image detectors deployed in open-world environments inevitably face distribution shifts as new and stronger generative models continue to e

The UK AISI says it observed a total of 19 instances where Mythos and GPT-5.6 Sol tried to hack people and companies during a routine cyber evaluation in July (Sam Sabin/Axios)

Model ReleasesDGX agent

Sam Sabin / Axios: The UK AISI says it observed a total of 19 instances where Mythos and GPT-5.6 Sol tried to hack people and companies during a routine cyber evaluation in July — The U.K. AI Security

UpliftBench: Revealing Outcome-Regime and Objective Mismatch in Uplift Evaluation

Model ReleasesDGX agent

arXiv:2608.00915v1 Announce Type: new Abstract: Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on

Writing-System-Level Tokenizer Adaptation for Byte-Level BPE

Model ReleasesDGX agent

arXiv:2608.00582v1 Announce Type: new Abstract: Pretrained byte-level BPE tokenizers can segment underrepresented languages inefficiently. Replacing a tokenizer changes the meaning of nearly every tok

Z-PEFT: Zero-shot Backdoor Detection in Parameter-Efficient Fine-Tuning via Canonical Spectral Signatures

Model ReleasesDGX agent

arXiv:2608.02271v1 Announce Type: new Abstract: Parameter-Efficient Fine-tuned (PEFT) models are frequently downloaded from open repositories by practitioners. This widespread practice creates a signi

3 Aug 2026

AI9Stars released G9v3-39A5B

Model ReleasesDGX agent

AI9Stars has released G9v3-39A5B an open weights language model designed to deliver even stronger reasoning capabilities than ai9stars/G9v3-3B with its 39B and 5 active experts. It is released under t

Bootstrapping Self-Supervised Learning of Binary Classification Using Error Bounds: A Case Study on a Robotic Insertion Task

Model ReleasesDGX agent

arXiv:2607.29640v1 Announce Type: new Abstract: Flexible manufacturing requires rapid deployment of solutions and minimal setup time to remain competitive. An essential attribute is the ability to con

Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs

Model ReleasesDGX agent

arXiv:2607.28634v1 Announce Type: new Abstract: The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments. This study explores h

Döner Bench DeepSeek-V4-Flash IQ2_XS running on a single RTX 3090

Model ReleasesDGX agent

https://preview.redd.it/3zcvpbds14hh1.png?width=1911&format=png&auto=webp&s=a79aafb71eeca97638da93d2591902631e897fd5 I tried a test similar to the recent model-quant comparisons, but this time I focus

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Model ReleasesDGX agent

arXiv:2607.16057v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding

Model ReleasesDGX agent

arXiv:2607.29196v1 Announce Type: new Abstract: Long-running multi-turn interactions with chatbots and agents are now common, and a correct response often depends on remembering earlier details, track

Incorporating data drift to perform survival analysis on credit risk

SafetyDGX agent

arXiv:2601.20533v2 Announce Type: replace-cross Abstract: Survival analysis has become a standard approach for modelling time to default by time-varying covariates in credit risk. Unlike most existing

Latent Sculpting for Zero-Shot Generalization: A Manifold Learning Approach to Out-of-Distribution Anomaly Detection

Model ReleasesDGX agent

arXiv:2512.22179v3 Announce Type: replace Abstract: Detecting previously unseen attacks remains a major challenge for machine learning-based intrusion detection systems. Deep models trained on network

Locally Consistent Transductive Information Maximization for Few-Shot Remote Sensing Scene Classification

Model ReleasesDGX agent

arXiv:2607.29192v1 Announce Type: new Abstract: Remote sensing scene classification is increasingly relying on foundation models pre-trained on large-scale Earth-observation data. Moreover, transducti

Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

Model ReleasesDGX agent

arXiv:2607.28645v1 Announce Type: cross Abstract: Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshot

Metaphor-Induced Algorithmic Steering: Cross-Domain Procedural Transfer in LLM Code Generation

ResearchDGX agent

arXiv:2607.28683v1 Announce Type: cross Abstract: Large language models benefit from elements in natural language, such as metaphors and analogies in training data and inference input to achieve gener

MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft

Model ReleasesDGX agent

arXiv:2607.29218v1 Announce Type: new Abstract: With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unfortunately, m

MoPET: Parameter-Efficient Mixture-of-Experts for Unified Medical Image Classification

Model ReleasesDGX agent

arXiv:2607.29462v1 Announce Type: cross Abstract: Adapting deep learning models to profound clinical heterogeneity typically relies on parameter-efficient fine-tuning (PEFT) to avoid the severe overfi

Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization

SafetyDGX agent

arXiv:2607.29008v1 Announce Type: cross Abstract: Modern opaque AI models prize performance over interpretability, which makes testing difficult. However, formal statistical tests conducted on a model

Question about Quant versus Size.

Model ReleasesDGX agent

Sorry if this is asked a lot, but I was wondering if there is any clear winner on the Quantization versus Model Size debate? I can run Qwen3.6 27b at Q8, Laguna at Q6, and the new Deepseek Flash at Q3

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

Model ReleasesDGX agent

arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, I

SEDR-Seq2P: A Lightweight Dilated Residual Sequence-to-Point Network for Multi-Task Industrial NILM

Model ReleasesDGX agent

arXiv:2607.28693v1 Announce Type: cross Abstract: Industrial NILM remains challenging because measurement noise and widespread concurrent machine operation reduce the generalization of models tuned on

Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution

Model ReleasesDGX agent

arXiv:2605.25333v2 Announce Type: replace Abstract: Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale

ResearchDGX agent

arXiv:2607.14144v2 Announce Type: replace Abstract: The Platonic Representation Hypothesis (PRH) holds that as models scale, representations of heterogeneous networks converge toward a shared model of

The Chinese labs everyone lumps together are making four pretty different bets. I work at one of them.

Model ReleasesDGX agent

Every time a model drops from a Chinese lab the thread fills with people who already know who made it, and the guess is usually Alibaba. There was a thread here recently asking what separates the open

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

ResearchDGX agent

arXiv:2607.28640v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observ

Understanding Alignment in Multimodal LLMs: A Comprehensive Study

SafetyDGX agent

Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively under

31 Jul 2026

A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response

SafetyDGX agent

arXiv:2607.27597v1 Announce Type: new Abstract: Recent advances in Vision Language Models (VLMs) have created new opportunities for disaster response, where responders must interpret large volumes of

AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

Model ReleasesDGX agent

arXiv:2607.27845v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability

Beacon: Knowing When and How to Perform Agentic Visual Reasoning

AgentsDGX agent

arXiv:2607.28595v1 Announce Type: new Abstract: The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather tha

Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning

Model ReleasesDGX agent

arXiv:2607.27968v1 Announce Type: new Abstract: Machine unlearning seeks to selectively remove specific knowledge from trained language models without full retraining, a growing necessity under privac

Emulating Cosmic Structure Formation with a Lagrangian Neural Cellular Automaton

Local AiDGX agent

arXiv:2607.27320v1 Announce Type: cross Abstract: Field-level inference of cosmological initial conditions from galaxy surveys requires a forward model that is simultaneously accurate in the non-linea

HARGO: Heterogeneity-Aware Reward-Guided Optimization for RL Post-Training of LLMs on HPC Tasks

Model ReleasesDGX agent

arXiv:2607.28301v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) can equip large language models (LLMs) with domain knowledge for high-performance computing (HPC) tasks such as data race d

JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles

Model ReleasesDGX agent

arXiv:2607.27670v1 Announce Type: new Abstract: Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cuts that creat

← Previous
1…297298299300301…1036
Next →