AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,428
  • Agents7,398
  • Applications5,301
  • Concepts5
  • Hardware1,785
  • Industry6,113
  • Local Ai4,833
  • Model Releases23,177
  • Research19,713
  • Safety13,092
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlog
86,428Total entries
1Added by human
86,427Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,016 results
Model Releases

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents

DGX agent

arXiv:2604.07429v1 Announce Type: new Abstract: Towards an embodied generalist for real-world interaction, Multimodal Large Language Model (MLLM) agents still suffer from challenging latency, sparse f

model-releasesarxiv-cs-cv
10 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

GEAR: GEometry-motion Alternating Refinement for Articulated Object Modeling with Gaussian Splatting

DGX agent

arXiv:2604.07728v1 Announce Type: new Abstract: High-fidelity interactive digital assets are essential for embodied intelligence and robotic interaction, yet articulated objects remain challenging to

researcharxiv-cs-cv
10 Apr 2026
Research

How Does Machine Learning Manage Complexity?

DGX agent

arXiv:2604.07233v1 Announce Type: new Abstract: We provide a computational complexity lens to understand the power of machine learning models, particularly their ability to model complex systems. Mach

researcharxiv-cs-lg
10 Apr 2026
Model Releases

IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures

DGX agent

arXiv:2604.07709v1 Announce Type: cross Abstract: Ask a frontier model how to taper six milligrams of alprazolam (psychiatrist retired, ten days of pills left, abrupt cessation causes seizures) and it

model-releasesarxiv-cs-cl
10 Apr 2026
Applications

Large Language Models for Outpatient Referral: Problem Definition, Benchmarking and Challenges

DGX agent

arXiv:2503.08292v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly applied to outpatient referral tasks across healthcare systems. However, there is a lack of stan

applicationsarxiv-cs-ai
10 Apr 2026
Model Releases

Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?

DGX agent

arXiv:2503.18018v2 Announce Type: replace Abstract: We demonstrate that large language models' (LLMs) mathematical reasoning is culturally sensitive: testing 14 models from Anthropic, OpenAI, Google,

model-releasesarxiv-cs-ai
10 Apr 2026
Research

LumiCtrl : Learning Illuminant Prompts for Lighting Control in Personalized Text-to-Image Models

DGX agent

arXiv:2512.17489v2 Announce Type: replace Abstract: Text-to-image (T2I) models have demonstrated remarkable progress in creative image generation, yet they still lack precise control over scene illumi

researcharxiv-cs-cv
10 Apr 2026
Safety

MDP modeling for multi-stage stochastic programs

DGX agent

arXiv:2509.22981v2 Announce Type: replace Abstract: We study a class of multi-stage stochastic programs, which incorporate modeling features from Markov decision processes (MDPs). This class includes

safetyarxiv-cs-lg
10 Apr 2026
Tutorials

OceanMAE: A Foundation Model for Ocean Remote Sensing

DGX agent

arXiv:2604.08171v1 Announce Type: new Abstract: Accurate ocean mapping is essential for applications such as bathymetry estimation, seabed characterization, marine litter detection, and ecosystem moni

tutorialsarxiv-cs-cv
10 Apr 2026
Safety

People in Washington get played, yet again We really should worry about cybersecurity - a lot – but Mythos is not the model these guys think…

DGX agent

People in Washington get played, yet again We really should worry about cybersecurity - a lot – but Mythos is not the model these guys think it is. (See my newsletter today for three reasons why it is

safetygary-marcus--x
10 Apr 2026
Research

SeLaR: Selective Latent Reasoning in Large Language Models

DGX agent

arXiv:2604.08299v1 Announce Type: new Abstract: Chain-of-Thought (CoT) has become a cornerstone of reasoning in large language models, yet its effectiveness is constrained by the limited expressivenes

researcharxiv-cs-cl
10 Apr 2026
Applications

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models

DGX agent

arXiv:2602.20231v2 Announce Type: replace-cross Abstract: Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-acti

applicationsarxiv-cs-cv
10 Apr 2026
Tools

We had Lin on stage: 'the future is millions of models — one per application, one per use case.' Jet delivered a masterclass on reinforcemen…

DGX agent

We had Lin on stage: 'the future is millions of models — one per application, one per use case.' Jet delivered a masterclass on reinforcement fine-tuning. Rob joined @WorkOS for some hot takes on the

toolsfireworks-ai--x
10 Apr 2026
Model Releases

maybe some nuance 😄 I don’t think anyone is “lying” about how great Mythos will be —> but there’s expectation misalignment between the Test…

DGX agent

maybe some nuance 😄 I don’t think anyone is “lying” about how great Mythos will be —> but there’s expectation misalignment between the Test Harness set up for Mythos and a belief it was given this cra

model-releasesharrison-chase--x
9 Apr 2026
Tools

Multimodal Embedding & Reranker Models with Sentence Transformers

DGX agent

The Sentence Transformers v5.4 update introduces first-class multimodal support, enabling the same familiar API to encode and compare texts, images, audio, and videos using both `SentenceTransforme...

toolshugging-face
9 Apr 2026
Industry

Refiant raises $5M to refine AI models with ‘nature-inspired’ energy efficiency

DGX agent

Artificial intelligence model compression startup Refiant AI said today it has raised $5 million in seed funding from VoLo Earth Ventures to try to put an end to the “arms race” that has ignited a mul

industrysiliconangle
9 Apr 2026
Tools

Pelicans for Meta's new Muse Spark models - plus I did a bit of a deep dive into the Code Interpreter and fascinating 'container.visual_grou…

DGX agent

Pelicans for Meta's new Muse Spark models - plus I did a bit of a deep dive into the Code Interpreter and fascinating 'container.visual_grounding' tools in their http://meta.ai chat UI https://simonwi

toolssimon-willison--x
8 Apr 2026
Safety

Want more proof that Anthropic's PR has no idea what it's talking about? The talk of Mythos being 'their most aligned model ever'. They coul…

DGX agent

Want more proof that Anthropic's PR has no idea what it's talking about? The talk of Mythos being 'their most aligned model ever'. They could perhaps truthfully speak about 'new high scores on our ali

safetyconnor-leahy--x
8 Apr 2026
Tools

Wrote up some thoughts on Anthropic's Project Glassing, where their latest Opus-beating model is available to partnered security research or…

DGX agent

Wrote up some thoughts on Anthropic's Project Glassing, where their latest Opus-beating model is available to partnered security research organizations only Given recent alarm bells raised by credible

toolssimon-willison--x
7 Apr 2026
Model Releases

Breaking the Compression Barrier: Cross-Architecture Compression Boundary Learning via Reverse Regrowth

DGX agent

arXiv:2608.16010v1 Announce Type: cross Abstract: Model compression is critical for deploying networks on resource-constrained edge devices. While pruning-based methods can significantly reduce model

model-releasesarxiv-cs-cv
18 Aug 2026
Model Releases

HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam

DGX agent

arXiv:2602.13964v4 Announce Type: replace Abstract: Humanity's Last Exam (HLE) has become a widely used benchmark for evaluating frontier large language models on challenging, multi-domain questions.

model-releasesarxiv-cs-cl
18 Aug 2026
Model Releases

OTel: Building Domain-Specialized Telecom LLM Foundations for Intelligent Networks

DGX agent

arXiv:2608.15436v1 Announce Type: new Abstract: Frontier AI models have advanced rapidly, but they still struggle with telecom-specific tasks. We present Open Telco (OTel), an open telecom AI resource

model-releasesarxiv-cs-ai
18 Aug 2026
Model Releases

Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies

DGX agent

arXiv:2608.16645v1 Announce Type: new Abstract: Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconst

model-releasesarxiv-cs-ai
18 Aug 2026
Model Releases

TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning

DGX agent

arXiv:2608.15055v1 Announce Type: new Abstract: Hypergraphs effectively model higher-order groupwise relationships beyond pairwise interactions, while pretrained language models (PLMs) and large langu

model-releasesarxiv-cs-ai
18 Aug 2026
Model Releases

VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding

DGX agent

arXiv:2608.14718v1 Announce Type: cross Abstract: Video understanding is a fundamental task for evaluating the capabilities of multimodal large language models (MLLMs). However, existing leading model

model-releasesarxiv-cs-cl
18 Aug 2026
Safety

AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations

DGX agent

arXiv:2608.14130v1 Announce Type: cross Abstract: Computer vision models for generated facial content, such as face editing and privacy protection, increasingly affect people, requiring similarity met

safetyarxiv-cs-ai
17 Aug 2026
Model Releases

CarbonBench: A Global Benchmark for Upscaling of Carbon Fluxes Using Zero-Shot Learning

DGX agent

arXiv:2603.09868v2 Announce Type: replace Abstract: Accurately quantifying terrestrial carbon exchange is essential for climate policy and carbon accounting, yet models must generalize to ecosystems u

model-releasesarxiv-cs-lg
17 Aug 2026
Model Releases

QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

DGX agent

arXiv:2608.13966v1 Announce Type: cross Abstract: As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aw

model-releasesarxiv-cs-cl
17 Aug 2026
Model Releases

SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

DGX agent

arXiv:2608.14277v1 Announce Type: cross Abstract: On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context r

model-releasesarxiv-cs-ai
17 Aug 2026
Model Releases

club-5060ti refresh: tested RTX 5060 Ti presets, a proper high-context harness, and Qwen3.8 27B

DGX agent

Quick update on the RTX 5060 Ti local LLM repo. It has changed quite a bit since my previous posts. The project started as a collection of practical notes and benchmark results. That was useful, but a

model-releasesr-localllama
15 Aug 2026
Model Releases

Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists

DGX agent

arXiv:2608.12345v1 Announce Type: new Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasur

model-releasesarxiv-cs-ai
14 Aug 2026
Model Releases

Qwen3.8-2.4T-A95B is live on Together AI. The Qwen team’s new open-weight MoE packs 2.4T total parameters with 95B active, a 1M context wind…

DGX agent

Qwen3.8-2.4T-A95B is live on Together AI. The Qwen team’s new open-weight MoE packs 2.4T total parameters with 95B active, a 1M context window, and strong coding & agent capabilities. Start building:

model-releasestogether-ai--x
14 Aug 2026
Model Releases

Steering the Language Axis: From Linear Decodability to Causal Control

DGX agent

arXiv:2608.12334v1 Announce Type: cross Abstract: Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly understood.

model-releasesarxiv-cs-ai
14 Aug 2026
Model Releases

TraVEL: Trajectory-Guided Video Embedding Learning for Driving-Video Retrieval

DGX agent

arXiv:2608.13495v1 Announce Type: new Abstract: Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured a

model-releasesarxiv-cs-cv
14 Aug 2026
Safety

Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed

DGX agent

arXiv:2608.11981v1 Announce Type: new Abstract: Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in r

safetyarxiv-cs-cl
13 Aug 2026
Model Releases

Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release

DGX agent

Qwen just released their first 3.8 model. The main addition in 3.8 is prompt-steered reasoning effort. You can tell the model how deeply to think by setting reasoning_effort to xhigh, medium, or low.

model-releasesr-localllama
13 Aug 2026
Model Releases

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

DGX agent

arXiv:2608.12036v1 Announce Type: new Abstract: AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain p

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques

DGX agent

arXiv:2608.11755v1 Announce Type: cross Abstract: Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards inc

model-releasesarxiv-cs-cl
13 Aug 2026
Model Releases

SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks

DGX agent

arXiv:2608.12127v1 Announce Type: new Abstract: Model routing aims to select the most suitable model from a candidate pool for each query, balancing quality and cost. Existing VLM routing research is

model-releasesarxiv-cs-cv
13 Aug 2026
Model Releases

The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance

DGX agent

arXiv:2608.11694v1 Announce Type: cross Abstract: A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same pr

model-releasesarxiv-cs-ai
13 Aug 2026
Model Releases

Variational Parameter Calibration with Physics-Aware Latent-Space Surrogates

DGX agent

arXiv:2608.11435v1 Announce Type: new Abstract: Forward and inverse modeling of parametric dynamical systems requires surrogate models that are not only accurate for state prediction, but also informa

model-releasesarxiv-cs-lg
13 Aug 2026
Agents

Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning

DGX agent

arXiv:2608.10438v1 Announce Type: new Abstract: Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outside world. Fo

agentsarxiv-cs-ai
12 Aug 2026
Model Releases

DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous Driving

DGX agent

arXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding

DGX agent

arXiv:2608.10764v1 Announce Type: new Abstract: Counterfactual video understanding evaluates whether models grasp physical and commonsense regularities. However, existing multiple-choice question (MCQ

model-releasesarxiv-cs-cv
12 Aug 2026
Model Releases

Hidden Reasoning from Claude and GPT are Decoded, and it is interesting

DGX agent

Yesteday a paper showed a gap that allows to see 100% of the reasoning tokens form ALL Claude and GPT models Stealing Reasoning Traces from Proprietary LLM APIs. check it out, they have published lots

model-releasesr-localllama
12 Aug 2026
Model Releases

Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability

DGX agent

arXiv:2608.10300v1 Announce Type: new Abstract: Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems, and human re

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

Rescene: band-limited stochastic forcing turns a frozen neural weather operator into a climate emulator

DGX agent

arXiv:2608.09971v1 Announce Type: cross Abstract: Over the past few years, the rapid development of machine learning (ML) models for weather forecasting has produced deterministic models whose medium-

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective Refinement

DGX agent

arXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models

model-releasesarxiv-cs-cv
12 Aug 2026
← Previous
1…248249250251252…1292
Next →