AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,597 results
20 Apr 2026

Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms

ResearchDGX agent

arXiv:2604.15842v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated impressive capabilities, yet their internal mechanisms for handling reasoning-intensive tasks remain unde

MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition

Model ReleasesDGX agent

arXiv:2604.16009v1 Announce Type: new Abstract: Metacognition, the ability to monitor and regulate one's own reasoning, remains under-evaluated in AI benchmarking. We introduce MEDLEY-BENCH, a benchma

QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.15859v1 Announce Type: cross Abstract: Forecasting has become a natural benchmark for reasoning under uncertainty. Yet existing evaluations of large language models remain limited to judgme

17 Apr 2026

Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes

ResearchDGX agent

arXiv:2604.14914v1 Announce Type: new Abstract: Text-driven inversion of generative models is a core paradigm for manipulating 2D or 3D content, unlocking numerous applications such as text-based edit

Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude. Powered by Claude Opus 4.7, our m…

Model ReleasesDGX agent

Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude. Powered by Claude Opus 4.7, our most capable vision model. Available in research preview on t

16 Apr 2026

C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions

ResearchDGX agent

arXiv:2604.13521v1 Announce Type: new Abstract: Neural network models with latent recurrent processing, where identical layers are recursively applied to the latent state, have gained attention as pro

Diagnostics for Individual-Level Prediction Instability in Machine Learning for Healthcare

ApplicationsDGX agent

arXiv:2603.00192v2 Announce Type: replace Abstract: In healthcare, predictive models increasingly inform patient-level decisions, yet little attention is paid to the variability in individual risk est

Structure- and Stability-Preserving Learning of Port-Hamiltonian Systems

ResearchDGX agent

arXiv:2604.13297v1 Announce Type: cross Abstract: This paper investigates the problem of data-driven modeling of port-Hamiltonian systems while preserving their intrinsic Hamiltonian structure and sta

15 Apr 2026

Accelerating Speculative Decoding with Block Diffusion Draft Trees

ResearchDGX agent

arXiv:2604.12989v1 Announce Type: new Abstract: Speculative decoding accelerates autoregressive language models by using a lightweight drafter to propose multiple future tokens, which the target model

Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss

Model ReleasesDGX agent

arXiv:2604.12911v1 Announce Type: cross Abstract: Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to p

SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents

Model ReleasesDGX agent

arXiv:2510.10073v2 Announce Type: replace-cross Abstract: Large vision-language model (LVLM)-based web agents are emerging as powerful tools for automating complex online tasks. However, when deployed

14 Apr 2026

A Mathematical Explanation of Transformers

ResearchDGX agent

arXiv:2510.03989v2 Announce Type: replace-cross Abstract: The Transformer architecture has revolutionized the field of sequence modeling and underpins the recent breakthroughs in large language models

AgencyBench: Benchmarking the Frontiers of Autonomous Agents in 1M-Token Real-World Contexts

Model ReleasesDGX agent

arXiv:2601.11044v3 Announce Type: replace Abstract: Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. Howev

Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation

TutorialsDGX agent

arXiv:2510.10925v2 Announce Type: replace-cross Abstract: Training student models on synthetic data generated by strong teacher models is a promising way to distilling the capabilities of teachers. Ho

GIANTS: Generative Insight Anticipation from Scientific Literature

Model ReleasesDGX agent

arXiv:2604.09793v1 Announce Type: cross Abstract: Scientific breakthroughs often emerge from synthesizing prior ideas into novel contributions. While language models (LMs) show promise in scientific d

Incentivizing Honesty among Competitors in Collaborative Learning and Optimization

Model ReleasesDGX agent

arXiv:2305.16272v5 Announce Type: replace Abstract: Collaborative learning techniques have the potential to enable training machine learning models that are superior to models trained on a single enti

Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

Model ReleasesDGX agent

arXiv:2604.11666v1 Announce Type: cross Abstract: As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dial

Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text

Model ReleasesDGX agent

arXiv:2601.17172v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly capable of generating personalized, persuasive text at scale, raising new questions about bias a

Why Don't You Know? Evaluating the Impact of Uncertainty Sources on Uncertainty Quantification in LLMs

ApplicationsDGX agent

arXiv:2604.10495v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed in real-world applications, reliable uncertainty quantification (UQ) becomes critical for safe

13 Apr 2026

A Compact Hybrid Convolution--Frequency State Space Network for Learned Image Compression

Model ReleasesDGX agent

arXiv:2511.20151v2 Announce Type: replace Abstract: Learned image compression (LIC) has recently benefited from Transformer- and state space models (SSM)- based backbones for modeling long-range depen

Adaptive Simulation Experiment for LLM Policy Optimization

Model ReleasesDGX agent

arXiv:2604.08779v1 Announce Type: new Abstract: Large language models (LLMs) have significant potential to improve operational efficiency in operations management. Deploying these models requires spec

AI Driven Soccer Analysis Using Computer Vision

ApplicationsDGX agent

arXiv:2604.08722v1 Announce Type: cross Abstract: Sport analysis is crucial for team performance since it provides actionable data that can inform coaching decisions, improve player performance, and e

BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation

ResearchDGX agent

arXiv:2604.09497v1 Announce Type: cross Abstract: Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse use cases.

Bias-Constrained Diffusion Schedules for PDE Emulations: Reconstruction Error Minimization and Efficient Unrolled Training

SafetyDGX agent

arXiv:2604.08357v2 Announce Type: replace Abstract: Conditional Diffusion Models are powerful surrogates for emulating complex spatiotemporal dynamics, yet they often fail to match the accuracy of det

On-the-Fly Adaptation to Quantization: Configuration-Aware LoRA for Efficient Fine-Tuning of Quantized LLMs

Model ReleasesDGX agent

arXiv:2509.25214v3 Announce Type: replace-cross Abstract: As increasingly large pre-trained models are released, deploying them on edge devices for privacy-preserving applications requires effective c

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos

Model ReleasesDGX agent

arXiv:2604.09037v1 Announce Type: cross Abstract: Current video benchmarks for multimodal large language models (MLLMs) focus on event recognition, temporal ordering, and long-context recall, but over

Task-Distributionally Robust Data-Free Meta-Learning

ApplicationsDGX agent

arXiv:2311.14756v2 Announce Type: replace-cross Abstract: Data-Free Meta-Learning (DFML) aims to enable efficient learning of unseen few-shot tasks, by meta-learning from multiple pre-trained models w

12 Apr 2026

After seeing that Claude Mythos marketing turned out to be, as expected, a scam, I wanted to make a master list of tricks being used to mark…

Model ReleasesDGX agent

After seeing that Claude Mythos marketing turned out to be, as expected, a scam, I wanted to make a master list of tricks being used to market LLMs. The master list includes statements directly from l

10 Apr 2026

Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection

Model ReleasesDGX agent

arXiv:2511.20162v2 Announce Type: replace Abstract: Large multi-modal models (LMMs) show increasing performance in realistic visual tasks for images and, more recently, for videos. For example, given

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting

Model ReleasesDGX agent

arXiv:2601.02670v2 Announce Type: replace Abstract: We introduce self-jailbreaking, a threat model in which an aligned LLM guides its own compromise. Unlike most jailbreak techniques, which oft

Calibration of a neural network ocean closure for improved mean state and variability

Model ReleasesDGX agent

arXiv:2604.06398v1 Announce Type: cross Abstract: Global ocean models exhibit biases in the mean state and variability, particularly at coarse resolution, where mesoscale eddies are unresolved. To add

Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment

Model ReleasesDGX agent

arXiv:2601.08258v3 Announce Type: replace Abstract: Large language models increasingly fail in a way that scalar accuracy cannot diagnose: they produce a sound reasoning trace and then abandon it unde

Dynamic Context Evolution for Scalable Synthetic Data Generation

Model ReleasesDGX agent

arXiv:2604.07147v1 Announce Type: cross Abstract: Large language models produce repetitive output when prompted independently across many batches, a phenomenon we term cross-batch mode collapse: the p

GLM-5.1 by @Zai_org is now #3 in Code Arena - surpassing Gemini 3.1 and GPT-5.4, and now on par with Claude Sonnet 4.6. The first frontier l…

Model ReleasesDGX agent

GLM-5.1 by @Zai_org is now #3 in Code Arena - surpassing Gemini 3.1 and GPT-5.4, and now on par with Claude Sonnet 4.6. The first frontier level open model to break into the top 3. It’s a major +90 po

JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency

Model ReleasesDGX agent

arXiv:2604.03044v2 Announce Type: replace-cross Abstract: We introduce JoyAI-LLM Flash, an efficient Mixture-of-Experts (MoE) language model designed to redefine the trade-off between strong performan

LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces

Model ReleasesDGX agent

arXiv:2604.06188v1 Announce Type: cross Abstract: People increasingly hold sustained, open-ended conversations with large language models (LLMs). Public reports and early studies suggest that, in such

LoFT: Parameter-Efficient Fine-Tuning for Long-tailed Semi-Supervised Learning in Open-World Scenarios

Model ReleasesDGX agent

arXiv:2509.09926v5 Announce Type: replace Abstract: Long-tailed semi-supervised learning (LTSSL) presents a formidable challenge where models must overcome the scarcity of tail samples while mitigatin

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale

Model ReleasesDGX agent

arXiv:2604.04771v2 Announce Type: replace-cross Abstract: Current document parsing methods advance primarily through model architecture innovation, while systematic engineering of training data remain

mlx @ollama!

Model ReleasesDGX agent

Ollama released a preview version (0.19) on March 31, 2026, built on top of Apple's open-source MLX framework, enabling local LLMs to run significantly faster on Apple Silicon Macs by leveraging th...

PLUME: Latent Reasoning Based Universal Multimodal Embedding

Model ReleasesDGX agent

arXiv:2604.02073v2 Announce Type: replace Abstract: Universal multimodal embedding (UME) maps heterogeneous inputs into a shared retrieval space with a single model. Recent approaches improve UME by g

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards

Model ReleasesDGX agent

arXiv:2512.01236v2 Announce Type: replace Abstract: Personalized generation models for a single subject have demonstrated remarkable effectiveness, highlighting their significant potential. However, w

Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition

Model ReleasesDGX agent

arXiv:2604.07884v1 Announce Type: new Abstract: High-fidelity generative models are increasingly needed in privacy-sensitive scenarios, where access to data is severely restricted due to regulatory an

SALLIE: Safeguarding Against Latent Language & Image Exploits

Model ReleasesDGX agent

arXiv:2604.06247v1 Announce Type: cross Abstract: Large Language Models (LLMs) and Vision-Language Models (VLMs) remain highly vulnerable to textual and visual jailbreaks, as well as prompt injections

T-Gated Adapter: A Lightweight Temporal Adapter for Vision-Language Medical Segmentation

Model ReleasesDGX agent

arXiv:2604.08167v1 Announce Type: new Abstract: Medical image segmentation traditionally relies on fully supervised 3D architectures that demand a large amount of dense, voxel-level annotations from c

TEMPER: Testing Emotional Perturbation in Quantitative Reasoning

Model ReleasesDGX agent

arXiv:2604.07801v1 Announce Type: new Abstract: Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, real-world quer

What’s new with Google Cloud

Model ReleasesDGX agent

Want to know the latest from Google Cloud? Find it here in one handy location. Check back regularly for our newest updates, announcements, resources, events, learning opportunities, and more. Tip: Not

9 Apr 2026

Gemma 4 31B brings dense multimodal reasoning to Together AI. Try Now: http://www.together.ai/models/gemma-4-31b

Model ReleasesDGX agent

Google's Gemma 4 31B is a dense multimodal model from Google DeepMind now available on Together AI's serverless infrastructure via the endpoint `google/gemma-4-31B-it`. It features a 256K context ...

Judging by my tl there is a growing gap in understanding of AI capability. The first issue I think is around recency and tier of use. I thin…

Model ReleasesDGX agent

Judging by my tl there is a growing gap in understanding of AI capability. The first issue I think is around recency and tier of use. I think a lot of people tried the free tier of ChatGPT somewhere l

8 Apr 2026

SWE-1.6 is lightning fast! Here's what 950 tok/s feels like - available in Windsurf today.

AgentsDGX agent

SWE-1.6 is lightning fast! Here's what 950 tok/s feels like - available in Windsurf today. Media We’re releasing SWE-1.6, our best model in both intelligence & model UX. SWE-1.6 matches our Preview mo

7 Apr 2026

Or try SWE-1.6 in Windsurf today: https://windsurf.com/

AgentsDGX agent

Cognition released SWE-1.6, their latest software engineering model optimized for both intelligence and 'model UX,' now generally available in the Windsurf IDE. It is free for the next three month...

14 Aug 2026

Federated Compositional Muon Optimizer for Matrix-Wise Models

ResearchDGX agent

arXiv:2608.12710v1 Announce Type: new Abstract: Muon, a more recently developed optimizer, is useful for matrix-wise models in AI areas. Although many works have studied Muon and its variants, these m

Large Language Models Persuade Without Planning Theory of Mind

SafetyDGX agent

arXiv:2602.17045v2 Announce Type: replace Abstract: A growing body of work attempts to evaluate the theory of mind (ToM) abilities of humans and large language models (LLMs) using static, non-interact

S2-HWM: Sparse Event-Structured Hierarchical World Model for Long-Horizon Surgical Robot Manipulation

ResearchDGX agent

arXiv:2608.13103v1 Announce Type: new Abstract: Long-horizon surgical robot manipulation is challenging because task rewards are sparse, while meaningful interaction changes occur at irregular interva

Splat-based Metal Artifact Reduction in Cone-Beam CT via Polychromatic Modeling

ApplicationsDGX agent

arXiv:2608.13159v1 Announce Type: new Abstract: Cone-beam computed tomography (CBCT) enables volumetric reconstruction from X-ray projections, but suffers from severe artifacts--especially beam harden

The AI Accountability Ecosystem in the Era of Language Models

ResearchDGX agent

arXiv:2608.12320v1 Announce Type: cross Abstract: This article reviews and updates the framework for accountability in AI based on account- ability ecosystems. We update the framework in light of the

Z.ai debuts GLM-5.3, using the same base model as GLM-5.2 with scaled post-training for stronger coding and cyber skills, with weights due in two weeks (Z.ai)

IndustryDGX agent

Z.ai: Z.ai debuts GLM-5.3, using the same base model as GLM-5.2 with scaled post-training for stronger coding and cyber skills, with weights due in two weeks — With GLM-5.2 we built the stack: IndexSh

13 Aug 2026

A consequence of how frontier models are trained is motivated reasoning, a phenomenon well studied in humans and discussed in this podcast f…

ApplicationsDGX agent

A consequence of how frontier models are trained is motivated reasoning, a phenomenon well studied in humans and discussed in this podcast from @PalisadeAI. Our thoughts, beliefs and reasonings tend t

Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines

HardwareDGX agent

arXiv:2608.11770v1 Announce Type: new Abstract: Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where

Transferable Above-Ground Biomass (AGB) Estimation Model from Multi-Sensor Data with Sparse Field Calibration

Local AiDGX agent

arXiv:2608.11638v1 Announce Type: cross Abstract: Spatially continuous quantification of forest above-ground biomass (AGB) is what makes carbon accounting credible and mitigation strategies actionable

12 Aug 2026

A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa

ApplicationsDGX agent

arXiv:2608.11053v1 Announce Type: cross Abstract: The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming. However, many e

← Previous
1…198199200201202…1010
Next →