AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,457
  • Agents7,399
  • Applications5,302
  • Concepts5
  • Hardware1,786
  • Industry6,117
  • Local Ai4,835
  • Model Releases23,193
  • Research19,715
  • Safety13,094
  • Syntheses17
  • Tools1,670
  • Tutorials3,324

Source
HumanDGX agent

86,457Total entries
1Added by human
86,456Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,039 results
10 Apr 2026

FinTruthQA: A Benchmark for AI-Driven Financial Disclosure Quality Assessment in Investor -- Firm Interactions

Model ReleasesDGX agent

arXiv:2406.12009v5 Announce Type: replace Abstract: Accurate and transparent financial information disclosure is essential for market efficiency, investor decision-making, and corporate governance. Ch

Got early access to a real-time interactive video model, here's what I found

Local AiDGX agent

I was unable to retrieve the specific Reddit post at the provided URL through my search. The post (reddit.com/r/StableDiffusion/comments/1shxmfk) did not surface in the search results, and I cannot...

@hwchase17 middleware was the right abstraction for it too. way more adoptable than asking everyone to restructure their agent setup

AgentsDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

LangChain's Middleware abstraction, introduced by Harrison Chase (@hwchase17) in LangChain 1.0 Alpha, addresses context engineering in AI agents by providing clean `before_model`, `after_model`, an...

Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling

SafetyDGX agent

arXiv:2604.07172v1 Announce Type: new Abstract: Calibration is central to reliable semantic uncertainty quantification, yet prior work has largely focused on discrimination, neglecting calibration. As

Luwen Technical Report

ApplicationsDGX agent

arXiv:2604.06737v1 Announce Type: cross Abstract: Large language models have demonstrated remarkable capabilities across a wide range of natural language processing tasks, yet their application in the

ParseBench: A Document Parsing Benchmark for AI Agents

Model ReleasesDGX agent

arXiv:2604.08538v1 Announce Type: new Abstract: AI agents are changing the requirements for document parsing. What matters is semantic correctness: parsed output must preserve the structure and

Physics-Informed Spectral Modeling for Hyperspectral Imaging

ResearchDGX agent

arXiv:2508.21618v2 Announce Type: replace-cross Abstract: We present PhISM, a physics-informed deep learning architecture that learns without supervision to explicitly disentangle hyperspectral observ

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models

Local AiDGX agent

arXiv:2604.06912v1 Announce Type: cross Abstract: MLLMs require high-resolution visual inputs for fine-grained tasks like document understanding and dense scene perception. However, current global res

Revisiting Radar Perception With Spectral Point Clouds

Model ReleasesDGX agent

arXiv:2604.08282v1 Announce Type: new Abstract: Radar perception models are trained with different inputs, from range-Doppler spectra to sparse point clouds. Dense spectra are assumed to outperform sp

See: https://open.substack.com/pub/garymarcus/p/three-reasons-to-think-that-the-claude?r=8tdk6&utm_medium=ios

Model ReleasesDGX agent

In an April 9, 2026 Substack post, AI skeptic Gary Marcus argues that Anthropic's announcement of its Claude 'Mythos' model was significantly overhyped, offering three key reasons for skepticism: t...

SPICE: Submodular Penalized Information-Conflict Selection for Efficient Large Language Model Training

ResearchDGX agent

arXiv:2601.23155v2 Announce Type: replace-cross Abstract: Information-based data selection for instruction tuning is compelling: maximizing the log-determinant of the Fisher information yields a monot

Synthetic Data for any Differentiable Target

SafetyDGX agent

arXiv:2604.08423v1 Announce Type: new Abstract: What are the limits of controlling language models via synthetic training data? We develop a reinforcement learning (RL) primitive, the Dataset Policy G

Through the Magnifying Glass: Adaptive Perception Magnification for Hallucination-Free VLM Decoding

ResearchDGX agent

arXiv:2503.10183v4 Announce Type: replace Abstract: Existing vision-language models (VLMs) often suffer from visual hallucination, where the generated responses contain inaccuracies that are not groun

TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories

Model ReleasesDGX agent

arXiv:2604.07223v1 Announce Type: cross Abstract: As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final outputs to int

Weakly Supervised Distillation of Hallucination Signals into Transformer Representations

Model ReleasesDGX agent

arXiv:2604.06277v1 Announce Type: new Abstract: Existing hallucination detection methods for large language models (LLMs) rely on external verification at inference time, requiring gold answers, retri

Weaves, Wires, and Morphisms: Formalizing and Implementing the Algebra of Deep Learning

ResearchDGX agent

arXiv:2604.07242v1 Announce Type: new Abstract: Despite deep learning models running well-defined mathematical functions, we lack a formal mathematical framework for describing model architectures. Ad

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

SafetyDGX agent

arXiv:2604.08524v1 Announce Type: cross Abstract: Applying steering vectors to large language models (LLMs) is an efficient and effective model alignment technique, but we lack an interpretable explan

9 Apr 2026

AI on the couch: Anthropic gives Claude 20 hours of psychiatry

Model ReleasesDGX agent

As part of the evaluation of its Claude Mythos model, Anthropic engaged a clinical psychiatrist for approximately 20 hours of evaluation sessions, describing Mythos as 'the most psychologically se...

With closed agent platforms like Claude Managed Agents, your agent's memory belongs to them, not you. It's locked behind their API. Agent in…

Model ReleasesDGX agent

With closed agent platforms like Claude Managed Agents, your agent's memory belongs to them, not you. It's locked behind their API. Agent infra should be open: open harness, open memory, model agnosti

8 Apr 2026

Today, our group at @Mila_Quebec and the lab of @francesarnold at @Caltech just released a new paper I contributed to, exploring how multimo…

Model ReleasesDGX agent

Today, our group at @Mila_Quebec and the lab of @francesarnold at @Caltech just released a new paper I contributed to, exploring how multimodal generative modeling could accelerate protein sciences! ⬇

7 Apr 2026

Frontier models will one shot just about anything few years Intelligence is compression Jevons law is more of a suggestion

IndustryDGX agent

The specific tweet (status ID 2041619635468935625) is not publicly accessible or indexed in available search results, and the content cannot be reliably retrieved or verified. I'm unable to produce...

GLM 5.1 is now LIVE in Atomic Chat SOTA for code & chat – now runs locally with TurboQuant Thanks to @zai_org for open-sourcing this frontie…

Model ReleasesDGX agent

GLM-5.1 is Z.ai's (zai-org) next-generation open-source flagship model for agentic engineering, achieving state-of-the-art performance on SWE-Bench Pro and significantly outperforming its predecess...

18 Aug 2026

A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation

ResearchDGX agent

arXiv:2608.16233v1 Announce Type: cross Abstract: Missing or degraded sequences can limit prostate multiparametric MRI. We developed MSCNet, a sequence-conditioned cross-modal generative framework for

A Deep Learning Model for Spatially Clustered Data via Differentiable Cluster Assignment

ResearchDGX agent

arXiv:2608.14968v1 Announce Type: cross Abstract: We consider nonparametric regression when the association between a response and its covariates changes across an unknown partition of a spatial domai

AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment

Model ReleasesDGX agent

arXiv:2608.16349v1 Announce Type: new Abstract: Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on stati

AeroGround: A Comprehensive Benchmark for Aerial-Ground Collaborative Reasoning

Model ReleasesDGX agent

arXiv:2608.14721v1 Announce Type: new Abstract: Vision-language models (VLMs) have been widely employed in understanding and reasoning tasks for unmanned aerial vehicles (UAVs). Existing UAV benchmark

Ask, Condition or Abstain: Reinforcement Learning for Missing-Premise Reasoning

Model ReleasesDGX agent

arXiv:2608.16554v1 Announce Type: new Abstract: Answer-only reinforcement learning (RL) trains reasoning models to solve fully specified problems, but many realistic queries omit a premise needed for

Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews

Model ReleasesDGX agent

arXiv:2608.14551v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for title-abstract screening in systematic reviews, but their decisions lack calibrated uncertainty.

Bye-bye, Bluebook? Automating Legal Drudgery With AI-Augmented Rule Following

Model ReleasesDGX agent

arXiv:2505.02763v2 Announce Type: replace-cross Abstract: One of the central promises of legal AI is to automate drudgery -- the formal, repetitive tasks of lawyers' work that consume time without cal

Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought

Model ReleasesDGX agent

arXiv:2603.18334v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) increasingly assist secure software development, their ability to meet the rigorous demands of Rust program ve

Disentangling Pictorial Cue Understanding from Language Bias in VLMs via Depth Ordering Task

Model ReleasesDGX agent

arXiv:2607.01503v2 Announce Type: replace Abstract: In this paper, we study depth perception of vision-language models (VLMs) to isolate the effects of pictorial depth cues and disentangle vision and

Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP

Model ReleasesDGX agent

arXiv:2608.14613v1 Announce Type: new Abstract: Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Go

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

SafetyDGX agent

arXiv:2608.16647v1 Announce Type: new Abstract: On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization be

From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving

Model ReleasesDGX agent

arXiv:2608.14771v1 Announce Type: new Abstract: Making language models solve constraint problems reliably often means having them translate the problem into a formal specification and delegating the s

Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning

Model ReleasesDGX agent

arXiv:2608.14804v1 Announce Type: new Abstract: Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, o

Is Ling 3 tiny underrated for its size?

Local AiDGX agent

I was checking out benchmarks of this model and apparantly better than Qwen3.5 9b reasoning across the bench on artificial analysis. I have used the 9b model for variety of stuff and it has been amazi

MoE Router-Guided Clustering for Heterogeneous Federated Instruction Tuning

Model ReleasesDGX agent

arXiv:2608.15311v1 Announce Type: new Abstract: Federated instruction fine-tuning enables Large Language Models (LLMs) to adapt to decentralized, privacy-sensitive data without requiring data sharing.

PixRestore: Unified Image Restoration via Pixel Diffusion Transformer

Model ReleasesDGX agent

arXiv:2608.16793v1 Announce Type: new Abstract: Unified image restoration (UIR) aims to recover high-quality (HQ) content from low-quality (LQ) images with different degradations using a single model.

Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters

Model ReleasesDGX agent

arXiv:2608.14792v1 Announce Type: cross Abstract: Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) behaviors in

Qwen 3.8 27b saved me $650+ in API costs this evening

Model ReleasesDGX agent

I've been experimenting with Qwen3.8-27B using DeepSeek Harness. It's a monster at long-horizon tasks, and the results were pretty wild. DeepSeek Harness ran on my Windows PC and connected over LAN to

The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning

Model ReleasesDGX agent

arXiv:2608.14558v1 Announce Type: new Abstract: Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstra

The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks

Model ReleasesDGX agent

arXiv:2608.16630v1 Announce Type: cross Abstract: Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context window. We mo

When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents

SafetyDGX agent

arXiv:2608.16806v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-step reasoning, and code generation, d

When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations

Model ReleasesDGX agent

arXiv:2608.15654v1 Announce Type: cross Abstract: Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-nat

17 Aug 2026

A Data-Driven Algorithm for Model-Free Control Synthesis

ResearchDGX agent

arXiv:2602.13157v2 Announce Type: replace-cross Abstract: Presented is an algorithm to synthesize the optimal infinite-horizon LQR feedback controller for continuous-time systems. The algorithm does n

A Graph-Based Reinforcement Learning Framework for Structured Drift Diagnosis and Recovery in Autonomous LLM Agents

Model ReleasesDGX agent

arXiv:2608.14109v1 Announce Type: new Abstract: Autonomous LLM agents are increasingly deployed in complex real-world workflows, yet they remain vulnerable to runtime behavioral drift, a silent deviat

Detecting Contaminated Code-Generation Prompt Batches via Influence Functions

Model ReleasesDGX agent

arXiv:2608.14303v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for code generation, yet they remain vulnerable to prompts that elicit insecure implementations. Exis

From single call to agents: five new Claude capabilities now available in Microsoft Foundry

Model ReleasesDGX agent

Structured outputs, web search, web fetch, MCP connector, and tool search are now available for Claude models hosted on Azure in Microsoft Foundry, turning a model endpoint into a production agent pla

How many tokens/second output are you getting with Qwen3.8-27B?

Model ReleasesDGX agent

Trying to get a feel for where I stand. If you can list your relevant hardware and model used, that would be awesome. Here's mine: Model: Qwen3.8-27B-heretic-ara, Q5_K_M GGUF T/s: ~30-32 t/sec (I thin

Non-Shattering at and Above the Dynamical Temperature in the Spherical Pure p-Spin Model

ResearchDGX agent

arXiv:2608.14369v1 Announce Type: cross Abstract: We consider the notion of shattering introduced by Ben Arous and Jagannath for spherical pure p-spin glasses with overlap q. For every pgeq 3 and 0sqr

NVIDIA Nemotron 3.5 Lightning now available in Amazon SageMaker JumpStart

Model ReleasesDGX agent

NVIDIA Nemotron 3.5 Lightning, an open model built for high-volume agentic workloads, is now available in Amazon SageMaker JumpStart. This post shows how to deploy the 30B Mixture-of-Experts model (3B

S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling

ResearchDGX agent

arXiv:2608.14029v1 Announce Type: new Abstract: Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual s

Some thoughts on Dario’s post: 1. Dario does not actually address Gavin Baker’s account of what he said – something he could easily deny if …

SafetyDGX agent

Some thoughts on Dario’s post: 1. Dario does not actually address Gavin Baker’s account of what he said – something he could easily deny if it were inaccurate. 2. Dario claims his critics live in a “b

16 Aug 2026

I am doing my best to stop back from this subject as it should be obvious what we are seeing here. However the arrogance that us commoners a…

Model ReleasesDGX agent

I am doing my best to stop back from this subject as it should be obvious what we are seeing here. However the arrogance that us commoners are too dumb to catch his grift must be addressed. Receipts:

Llama.cpp ROCm 7.2->7.14 upgrade, Radeon 780m iGPU benchmarks: ROCm vs Vulkan

Model ReleasesDGX agent

With all the new models released recently one important upgrade went unnoticed: Llama.cpp bumped ROCm from 7.2 to 7.14. I was waiting for that because in 7.14 support for gfx1103 (Radeon 780m) was int

Newer commits removed the Qwen 35B

Model ReleasesDGX agent

In this commits, the 35B model was removed. Looks like it's confirming the 35B model won't get released. I think they need to be made aware how big the 35 moe is widely used. Think need to make noise

15 Aug 2026

Anthropic details Claude's text watermark: it only shows Claude was likely involved, is sparse in code and factual text, and disappears after a full rewrite (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic details Claude's text watermark: it only shows Claude was likely involved, is sparse in code and factual text, and disappears after a full rewrite — Future Claude models will gene

14 Aug 2026

Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence

Model ReleasesDGX agent

arXiv:2608.12895v1 Announce Type: new Abstract: Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that

Learning Under Treatment-Induced Label Indeterminacy with Expert Annotations of Counterfactual Outcomes: A Case Study in Neurological Prognostication

SafetyDGX agent

arXiv:2608.12477v1 Announce Type: new Abstract: Clinical prediction models are often developed as if the outcome of interest were cleanly observed for every patient. This assumption fails when treatme

Qwen Live EP2 - Agent First: Multimodal Gets to Work - Streamed live 11 hours ago

Model ReleasesDGX agent

Possibly best way to spend current hour before grabbing 27B model(Instead of opening duplicate repeated 27B posts here). I tried to grab summary(thought of including in this thread) of this video usin

← Previous
1…268269270271272…1034
Next →