AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlog
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,553 results
Model Releases

Access Sets Matter: Budgeting Expert Reads for Scalable Weight-Space Model Merging

DGX agent

arXiv:2605.29489v1 Announce Type: new Abstract: Weight-space model merging is usually formulated as an algebraic operation on checkpoints, yet at LLM scale the limiting resource is often the set of ex

model-releasesarxiv-cs-lg
29 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Auditing Training Data in Generative Music Models via Black-Box Membership Inference

DGX agent

arXiv:2605.29202v1 Announce Type: new Abstract: Recent advances in text-to-music generation enable high-fidelity synthesis of structured musical audio, raising growing concerns about data provenance,

safetyarxiv-cs-lg
29 May 2026
Model Releases

Benchmarking Positional Encoding Strategies for Transformer-Based EEG Foundation Models

DGX agent

arXiv:2605.29754v1 Announce Type: new Abstract: Electroencephalography (EEG) is a widely used non-invasive technique for measuring brain activity in brain-computer interface (BCI) applications. Superv

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

DGX agent

arXiv:2605.30162v1 Announce Type: new Abstract: Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model

model-releasesarxiv-cs-ai
29 May 2026
Safety

DLM-SWAI: Steering Diffusion Language Models Before They Unmask

DGX agent

arXiv:2605.29626v1 Announce Type: cross Abstract: Steering language model generation toward desired textual properties is essential for practical deployment, and inference-time methods are particularl

safetyarxiv-cs-ai
29 May 2026
Research

Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding

DGX agent

arXiv:2605.29707v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference by drafting multiple tokens and verifying them in parallel with the target model. However, its practical

researcharxiv-cs-cl
29 May 2026
Tutorials

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models

DGX agent

arXiv:2605.29438v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are a powerful paradigm for generalist robotic control. However, their high computational cost and limited control f

tutorialsarxiv-cs-ro
29 May 2026
Research

Large-Scale AI and Foundation Models for Neuroscience: A Comprehensive Review

DGX agent

arXiv:2510.16658v3 Announce Type: replace Abstract: The development of large-scale artificial intelligence (AI) models is influencing neuroscience research by enabling end-to-end learning from raw bra

researcharxiv-cs-ai
29 May 2026
Research

LLMSurgeon: Diagnosing Data Mixture of Large Language Models

DGX agent

arXiv:2605.30348v1 Announce Type: cross Abstract: The pretraining data mixture of Large Language Models (LLMs) constitutes their 'digital DNA', shaping model behaviors, capabilities, and failure modes

researcharxiv-cs-ai
29 May 2026
Research

Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context

DGX agent

arXiv:2510.06182v2 Announce Type: replace Abstract: A key component of in-context reasoning is the ability of language models (LMs) to bind entities for later retrieval. For example, an LM might repre

researcharxiv-cs-cl
29 May 2026
Research

Mixing Vector Model for Copolymer Inference via Mixed Integer Linear Programming

DGX agent

arXiv:2605.29329v1 Announce Type: cross Abstract: A novel two-phase molecule inference framework, mol-infer, has recently been developed to infer chemical graphs with prescribed abstract structures an

researcharxiv-cs-lg
29 May 2026
Model Releases

models underestimate how much work it takes (token usage) to accomplish a task, just like us

DGX agent

models underestimate how much work it takes (token usage) to accomplish a task, just like us 🧵 Claude-Opus-4.8 takes you too much tokens - but is this issue general across agents? Do agents know how m

model-releasesyohei-nakajima--x
29 May 2026
Tutorials

Multi-level Collaborative Distillation Meets Global Workspace Model: A Unified Framework for OCIL

DGX agent

arXiv:2508.08677v2 Announce Type: replace-cross Abstract: Online Class-Incremental Learning (OCIL) enables models to learn continuously from non-i.i.d. data streams. Since samples of the data streams

tutorialsarxiv-cs-cv
29 May 2026
Local Ai

New LFM2.5 8b A1b model!!

DGX agent

Liquid AI released LFM2.5-8B-A1B, a device-optimized model designed to power real-life applications on phones, laptops, PCs, robots, and lightweight server-side use-cases. The model is a fast, memory-

local-air-ollama
29 May 2026
Model Releases

OOD-GraphLLM: Graph Large Language Model for Out-of-Distribution Generalized Drug Synergy Prediction

DGX agent

arXiv:2605.30247v1 Announce Type: new Abstract: Drug synergy prediction (DSP) aims to identify efficacious drug combinations under various cellular contexts with different targets. However, the contin

model-releasesarxiv-cs-lg
29 May 2026
Applications

Order-Agnostic Autoregressive Modelling with Missing Data

DGX agent

arXiv:2605.06355v2 Announce Type: replace Abstract: Order-Agnostic autoregressive models have demonstrated strong performance in deep generative modeling, yet their use in settings with incomplete dat

applicationsarxiv-cs-lg
29 May 2026
Model Releases

Orthogonal Concept Erasure for Diffusion Models

DGX agent

arXiv:2605.28902v1 Announce Type: new Abstract: Concept erasure has emerged as a promising approach to mitigate undesired or unsafe content in diffusion models, yet existing methods still face signifi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Parallax: Parameterized Local Linear Attention for Language Modeling

DGX agent

arXiv:2605.29157v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become the central paradigm in artificial intelligence, yet the core computational primitive of attention has remain

model-releasesarxiv-cs-ai
29 May 2026
Applications

PassNet: Scaling Large Language Models for Graph Compiler Pass Generation

DGX agent

arXiv:2605.29357v1 Announce Type: new Abstract: Modern tensor compilers such as TorchInductor deliver substantial speedups on mainstream models, yet face a systematic performance ceiling on long-tail

applicationsarxiv-cs-ai
29 May 2026
Research

Rare Event Analysis of Large Language Models

DGX agent

arXiv:2602.06791v2 Announce Type: replace Abstract: Being probabilistic models, during inference large language models (LLMs) display rare events: behaviour that is far from typical but highly signifi

researcharxiv-cs-lg
29 May 2026
Model Releases

Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

DGX agent

arXiv:2605.29327v1 Announce Type: new Abstract: Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high trai

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling

DGX agent

arXiv:2602.11065v2 Announce Type: replace-cross Abstract: Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing thi

model-releasesarxiv-cs-ai
29 May 2026
Research

SCoOP: Semantic Consistent Opinion Pooling for Uncertainty Quantification in Multiple Vision-Language Model Systems

DGX agent

arXiv:2603.23853v3 Announce Type: replace Abstract: Combining multiple Vision-Language Models (VLMs) can enhance multimodal reasoning and robustness, but aggregating heterogeneous models' outputs ampl

researcharxiv-cs-ai
29 May 2026
Model Releases

Sequential Physics-Constrained Neural Operator Forward Modeling for the extit{Norne} Reservoir System

DGX agent

arXiv:2605.28909v1 Announce Type: new Abstract: We develop a comprehensive mathematical and computational framework for sequential surrogate modeling of three-phase black-oil reservoir dynamics using

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

SMolLM: Small Language Models Learn Small Molecular Grammar

DGX agent

arXiv:2605.06322v2 Announce Type: replace Abstract: Language models for molecular design have scaled to hundreds of millions of parameters, yet how they learn chemical grammar is poorly understood. We

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning

DGX agent

arXiv:2605.30257v1 Announce Type: new Abstract: We present Stable-Layers, a reinforcement learning framework that eliminates the need for paired supervision by fine-tuning a pretrained layer decomposi

model-releasesarxiv-cs-cv
29 May 2026
Research

Steering Language Models Before They Speak: Logit-Level Interventions

DGX agent

arXiv:2601.10960v2 Announce Type: replace-cross Abstract: Controllable generation requires language models to realize output characteristics such as reading level, politeness, and toxicity. Existing s

researcharxiv-cs-ai
29 May 2026
Research

Tailoring the Curriculum: Student-Centered Reasoning Distillation via Dynamic Data-Model Compatibility

DGX agent

arXiv:2605.29229v1 Announce Type: new Abstract: Reasoning distillation transfers complex reasoning abilities from large language models (LLMs) to smaller ones, yet its success depends on how well the

researcharxiv-cs-ai
29 May 2026
Research

Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usage

DGX agent

arXiv:2605.30040v1 Announce Type: cross Abstract: Per-token billing is now the standard pricing model for commercial large language models (LLMs), so the honesty of reported token counts directly affe

researcharxiv-cs-ai
29 May 2026
Research

TrojanTO: Action-Level Backdoor Attacks against Trajectory Optimization Models

DGX agent

arXiv:2506.12815v2 Announce Type: replace Abstract: Recent advances in Trajectory Optimization (TO) models have achieved remarkable success in offline reinforcement learning. However, their vulnerabil

researcharxiv-cs-lg
29 May 2026
Model Releases

AI researchers ran 15-day simulations of worlds governed by different AI models: Claude Sonnet 4.6 recorded no crimes, while Gemini 3 Flash had the most at 683 (Jake Angelo/Fortune)

DGX agent

Jake Angelo / Fortune: AI researchers ran 15-day simulations of worlds governed by different AI models: Claude Sonnet 4.6 recorded no crimes, while Gemini 3 Flash had the most at 683 — Imagine a world

model-releasestechmeme
28 May 2026
Applications

Analyzing Cancer Patients' Experiences with Embedding-based Topic Modeling and LLMs

DGX agent

arXiv:2601.12154v2 Announce Type: replace Abstract: This study investigates the use of neural topic modeling and LLMs to uncover meaningful themes from patient storytelling data, to offer insights tha

applicationsarxiv-cs-cl
28 May 2026
Model Releases

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring

DGX agent

arXiv:2502.05242v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are becoming increasingly capable, but the mechanisms of their thinking and decision-making processes remain uncl

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models

DGX agent

arXiv:2605.27383v1 Announce Type: cross Abstract: Spoken Language Models (SLMs) have emerged as a promising paradigm for speech synthesis by bypassing explicit grapheme-to-phoneme pipelines. However,

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Claude Opus 4.8 is out today. It's our strongest coding model yet: up on SWE-bench Pro (from 64.3 to 69.2) and noticeably more honest about …

DGX agent

Claude Opus 4.8 is out today. It's our strongest coding model yet: up on SWE-bench Pro (from 64.3 to 69.2) and noticeably more honest about its own work. It tells you when it's unsure and catches its

model-releasesboris-cherny--x
28 May 2026
Model Releases

Colosseum V2: Benchmarking Generalization for Vision Language Action Models

DGX agent

arXiv:2605.27759v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language

model-releasesarxiv-cs-ro
28 May 2026
Research

Entropy Distribution as a Fingerprint for Hallucinations in Generative Models

DGX agent

arXiv:2605.28264v1 Announce Type: new Abstract: Large Language Models (LLMs) often generate factually incorrect outputs, commonly termed hallucinations, that undermine trust and limit deployment in hi

researcharxiv-cs-ai
28 May 2026
Research

FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models

DGX agent

arXiv:2512.16483v2 Announce Type: replace Abstract: Visual Autoregressive (VAR) modeling departs from the next-token prediction paradigm of traditional Autoregressive (AR) models through next-scale pr

researcharxiv-cs-cv
28 May 2026
Model Releases

IBM and Red Hat commit $5B to establish a new model for open-source software, dubbed Project Lightwell, and will deploy 20,000 engineers, supported by AI (Connor Hart/Wall Street Journal)

DGX agent

Connor Hart / Wall Street Journal: IBM and Red Hat commit $5B to establish a new model for open-source software, dubbed Project Lightwell, and will deploy 20,000 engineers, supported by AI — Project L

model-releasestechmeme
28 May 2026
Research

Measuring Form and Function in Language Models

DGX agent

arXiv:2605.28616v1 Announce Type: cross Abstract: We introduce quantitative metrics for child language acquisition to evaluate language models. Our focus is on the formal syntactic and functional disc

researcharxiv-cs-ai
28 May 2026
Model Releases

MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models

DGX agent

arXiv:2605.28009v1 Announce Type: cross Abstract: Memory-augmented large language models extend reasoning beyond a fixed context window by maintaining long-term memory across interactions. However, ex

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems

DGX agent

arXiv:2605.28732v1 Announce Type: cross Abstract: Memory is essential for enabling large language models to support long-horizon reasoning, yet existing memory systems remain unreliable and difficult

model-releasesarxiv-cs-ai
28 May 2026
Industry

OpenCode Now Supports DigitalOcean Inference Router for Intelligent Model Routing

DGX agent

OpenCode now integrates with DigitalOcean's Inference Router, enabling intelligent routing of AI model requests across distributed infrastructure. This integration allows developers to optimize model

industrydigitalocean
28 May 2026
Model Releases

Particle-Guided Diffusion Models for Partial Differential Equations

DGX agent

arXiv:2601.23262v2 Announce Type: replace Abstract: We introduce a guided stochastic sampling method that augments sampling from diffusion models with physics-based guidance derived from partial diffe

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models

DGX agent

arXiv:2503.01829v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) demonstrate persuasive capabilities that rival human-level persuasion. While these capabilities can be used for s

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Pittsburgh-based Gray Swan, which stress-tests AI models for top frontier AI labs, raised a 40M Series A at a 200M valuation co-led by Wing VC and Madrona (Rashi Shrivastava/Forbes)

DGX agent

Rashi Shrivastava / Forbes: Pittsburgh-based Gray Swan, which stress-tests AI models for top frontier AI labs, raised a 40M Series A at a 200M valuation co-led by Wing VC and Madrona — Gray Swan works

model-releasestechmeme
28 May 2026
Model Releases

Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents

DGX agent

arXiv:2605.28201v1 Announce Type: new Abstract: Large Language Model (LLM) agents remain vulnerable to safety threats from the external environment, where attackers inject adversarial content into ext

model-releasesarxiv-cs-ai
28 May 2026
Research

Quantum Machine Learning-based 6G edge Network: Enabling Adaptive Communication and Model Aggregation

DGX agent

arXiv:2605.27417v1 Announce Type: cross Abstract: With the advent of sixth-generation (6G) mobile communication technology, vehicle-to-everything (V2X) communication faces unprecedented challenges in

researcharxiv-cs-ai
28 May 2026
← Previous
1…137138139140141…1262
Next →