AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,515 results
12 May 2026

Attention Drift: What Autoregressive Speculative Decoding Models Learn

TutorialsDGX agent

arXiv:2605.09992v1 Announce Type: cross Abstract: Speculative decoding accelerates LLM inference by drafting future tokens with a small model, but drafter models degrade sharply under template perturb

Bi-CoG: Bi-Consistency-Guided Self-Training for Vision-Language Models

SafetyDGX agent

arXiv:2510.20477v2 Announce Type: replace Abstract: Exploiting unlabeled data through semi-supervised learning (SSL) or leveraging pre-trained models via fine-tuning are two prevailing paradigms for a

C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.10744v1 Announce Type: new Abstract: Safety-critical planning in complex environments, particularly at urban intersections, remains a fundamental challenge for autonomous driving. Existing

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

DARE: Diffusion Language Model Activation Reuse for Efficient Inference

ResearchDGX agent

arXiv:2605.08134v1 Announce Type: cross Abstract: Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to auto-regressive (AR) models, offering greater expressive capacity a

Do multimodal models imagine electric sheep?

ResearchDGX agent

arXiv:2605.09693v1 Announce Type: cross Abstract: Yes. We find that large multimodal models develop mental imagery when solving spatial puzzles, and they do imagine sheep when solving sheep puzzles. W

DriveFuture: Future-Aware Latent World Models for Autonomous Driving

AgentsDGX agent

arXiv:2605.09701v1 Announce Type: new Abstract: Existing latent world models for autonomous driving have opened a promising path toward future-aware driving intelligence. However, they typically treat

Dystruct: Dynamically Structured Diffusion Language Model Decoding via Bayesian Inference

Local AiDGX agent

arXiv:2605.09820v1 Announce Type: new Abstract: Diffusion language models (DLMs) have recently emerged as a promising alternative to autoregressive models, primarily due to their ability to enable par

Edit-Based Refinement for Parallel Masked Diffusion Language Models

ResearchDGX agent

arXiv:2605.09603v1 Announce Type: new Abstract: Masked diffusion language models enable parallel token generation and offer improved decoding efficiency over autoregressive models. However, their perf

Exploitation Without Deception: Dark Triad Feature Steering Reveals Separable Antisocial Circuits in Language Models

Model ReleasesDGX agent

arXiv:2605.09773v1 Announce Type: cross Abstract: We use sparse autoencoder (SAE) feature steering to amplify Dark Triad personality traits (Machiavellianism, narcissism, and psychopathy) in Llama-3.3

Failing Forward: Adaptive Failure-Informed Learning for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.08434v1 Announce Type: new Abstract: Vision-language-action (VLA) models provide a promising paradigm for scalable robotic manipulation, yet their reliance on success-only behavioral clonin

FedGMI: Generative Model-Driven Federated Learning for Probabilistic Mixture Inference

Local AiDGX agent

arXiv:2605.08760v1 Announce Type: new Abstract: Federated Learning (FL) facilitates collaborative model training across decentralized clients while preserving data privacy by avoiding raw data exchang

Filtering Memorization from Parameter-Space in Diffusion Models

Model ReleasesDGX agent

arXiv:2605.10439v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become a widely used mechanism for customizing diffusion models, enabling users to inject new visual concepts or styles t

Grounding the Score: Explicit Visual Premise Verification for Reliable Vision-Language Process Reward Models

Model ReleasesDGX agent

arXiv:2603.16253v2 Announce Type: replace-cross Abstract: Vision-language process reward models (VL-PRMs) are increasingly used to score intermediate reasoning steps and rerank candidates under test-t

High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models

ResearchDGX agent

arXiv:2512.21815v2 Announce Type: replace Abstract: Vision-language models (VLMs) achieve remarkable performance but remain vulnerable to adversarial attacks. Entropy, as a measure of model uncertaint

Illusion-Aware Visual Preprocessing and Anti-Illusion Prompting for Classic Illusion Understanding in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.08841v1 Announce Type: new Abstract: Vision-Language Models (VLMs) exhibit systematic bias toward visual illusions, recalling memorized facts rather than perceiving actual visual difference

Learngene Search Across Multiple Datasets for Building Variable-Sized Models

ResearchDGX agent

arXiv:2605.08209v1 Announce Type: new Abstract: Deep learning methods are widely used under diverse resource constraints, resulting in models of varying sizes, such as the Vision Transformer (ViT) ser

Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models

Model ReleasesDGX agent

arXiv:2510.08592v3 Announce Type: replace-cross Abstract: Test-Time Scaling (TTS) improves LLM reasoning by exploring multiple candidate responses and then operating over this set to find the best out

LiteParse is the best open-source, model-free document parser for AI agents. Run it over over 50+ document types, and it will parse dense pa…

Model ReleasesDGX agent

LiteParse is the best open-source, model-free document parser for AI agents. Run it over over 50+ document types, and it will parse dense pages with complex text layouts and tables, and it will extrac

LLiMba: Sardinian on a Single GPU -- Adapting a 3B Language Model to a Vanishing Romance Language

Model ReleasesDGX agent

arXiv:2605.09015v1 Announce Type: new Abstract: Sardinian, a Romance language with roughly one million speakers, has minimal presence in modern NLP. Commercial services do not support it, and current

Med-StepBench: A Hierarchical Reasoning Framework for Evaluating Hallucinations in Medical Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.10002v1 Announce Type: new Abstract: Large vision-language models (VLMs) demonstrate strong performance in medical image understanding, but frequently generate clinically plausible yet inco

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology

Model ReleasesDGX agent

arXiv:2605.09152v1 Announce Type: new Abstract: Deciphering animal intent is a fundamental challenge in computational ethology, largely because of semantic aliasing, the phenomenon where identical ext

On the Generation and Mitigation of Harmful Geometry in Image-to-3D Models

SafetyDGX agent

arXiv:2605.09606v1 Announce Type: cross Abstract: Recent advances in image-to-3D models have significantly improved the fidelity and accessibility of 3D content creation. Such a powerful reconstructio

Pretraining large language models with MXFP4

Model ReleasesDGX agent

arXiv:2605.09825v1 Announce Type: cross Abstract: Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable? We a

SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models

SafetyDGX agent

arXiv:2510.20129v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, where adversarially crafted prompts induce policy-violating responses des

SLayerGen: a Crystal Generative Model for all Space and Layer Groups

ResearchDGX agent

arXiv:2605.08262v1 Announce Type: cross Abstract: Crystal generative models have shown rapid progress for accelerating the discovery of bulk, periodic materials. However, many material systems such as

SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training

ResearchDGX agent

arXiv:2605.08738v1 Announce Type: cross Abstract: Structured pruning and knowledge distillation (KD) are typical techniques for compressing large language models, but it remains unclear how they shoul

This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization,…

Model ReleasesDGX agent

This NVIDIA remains the strongest platform for large-model inference at scale. Prefill/decode disaggregation, Blackwell-native quantization, custom kernels, and rack-scale NVLink turn GB200 into faste

Towards a Large Language-Vision Question Answering Model for MSTAR Automatic Target Recognition

Model ReleasesDGX agent

arXiv:2605.10772v1 Announce Type: cross Abstract: Large language-vision models (LLVM), such as OpenAI's ChatGPT and GPT-4, have gained prominence as powerful tools for analyzing text and imagery. The

Towards Backdoor-Based Ownership Verification for Vision-Language-Action Models

ResearchDGX agent

arXiv:2605.09005v1 Announce Type: cross Abstract: Vision-Language-Action models (VLAs) support generalist robotic control by enabling end-to-end decision policies directly from multi-modal inputs. As

Towards Compact Sign Language Translation: Frame Rate and Model Size Trade-offs

Model ReleasesDGX agent

arXiv:2605.09554v1 Announce Type: new Abstract: Sign Language Translation (SLT) converts sign language videos into spoken-language text, bridging communication between Deaf and hearing communities. Cu

Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models

Model ReleasesDGX agent

arXiv:2605.08974v1 Announce Type: cross Abstract: While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We arg

Uncovering Intra-expert Activation Sparsity for Efficient Mixture-of-Expert Model Execution

ResearchDGX agent

arXiv:2605.08575v1 Announce Type: cross Abstract: Mixture of Experts (MoE) architecture has become the standard for state-of-the-art large language models, owing to its computational efficiency throug

Understanding Asynchronous Inference Methods for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.08168v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models offer a promising path to generalist robot control, but their inference latency causes observation staleness when

Unsupervised Process Reward Models

SafetyDGX agent

arXiv:2605.10158v1 Announce Type: new Abstract: Process Reward Models (PRMs) are a powerful mechanism for steering large language model reasoning by providing fine-grained, step-level supervision. How

11 May 2026

Amortized Multi-Objective Optimization Across Tasks with Generative Solution Modeling

Model ReleasesDGX agent

arXiv:2511.09598v5 Announce Type: replace Abstract: Many real-world applications require solving families of expensive multi-objective optimization problems~(EMOPs) under varying operational condition

BGM-IV: an AI-powered Bayesian generative modeling approach for instrumental variable analysis

Model ReleasesDGX agent

arXiv:2605.07029v1 Announce Type: cross Abstract: Instrumental-variable (IV) regression enables causal estimation under endogeneity, but modern IV problems often involve nonlinear structural effects a

DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization

Model ReleasesDGX agent

arXiv:2512.14263v2 Announce Type: replace-cross Abstract: Preferential Bayesian Optimization (PBO) aims to find a decision-maker's most preferred solution in as few pairwise comparisons as possible. E

Efficient Data Selection for Multimodal Models via Incremental Optimization Utility

Model ReleasesDGX agent

arXiv:2605.07488v1 Announce Type: new Abstract: The scaling of Large Multimodal Models (LMMs) is constrained by the quality-quantity trade-off inherent in synthetic data. Previous approaches, such as

Flow-OPD: On-Policy Distillation for Flow Matching Models

SafetyDGX agent

arXiv:2605.08063v1 Announce Type: cross Abstract: Existing Flow Matching (FM) text-to-image models suffer from two critical bottlenecks under multi-task alignment: the reward sparsity induced by scala

Frequency-Aware Model Parameter Explorer: A new attribution method for improving explainability

Model ReleasesDGX agent

arXiv:2510.03245v2 Announce Type: replace-cross Abstract: State-of-the-art attribution methods rely on adversarial sample generation that applies an all-pass filter across the frequency spectrum, disc

Introducing Daybreak: frontier AI for cyber defenders. Daybreak brings together the most capable OpenAI models, Codex, and our security part…

Model ReleasesDGX agent

Introducing Daybreak: frontier AI for cyber defenders. Daybreak brings together the most capable OpenAI models, Codex, and our security partners to accelerate cyber defense and continuously secure sof

It is one reason why I think the push for smaller, local models is more complicated than people think. If you want good answers, especially …

ApplicationsDGX agent

It is one reason why I think the push for smaller, local models is more complicated than people think. If you want good answers, especially good answers to unexpected problems, frontier models will ge

MicroBi-ConvLSTM: An Ultra-Lightweight Efficient Model for Human Activity Recognition on Resource Constrained Devices

Model ReleasesDGX agent

arXiv:2602.06523v2 Announce Type: replace Abstract: Human Activity Recognition (HAR) on resource constrained wearables requires models that balance accuracy against strict memory and computational bud

NoiseGate: Learning Per-Latent Timestep Schedules as Information Gating in World Action Models

SafetyDGX agent

arXiv:2605.07794v1 Announce Type: new Abstract: World Action Models (WAMs) are an emerging family of policies that tie robot action generation to future-observation modeling. In this work, we focus on

Normalizing Trajectory Models

ResearchDGX agent

arXiv:2605.08078v1 Announce Type: new Abstract: Diffusion-based models decompose sampling into many small Gaussian denoising steps -- an assumption that breaks down when generation is compressed to a

Our research, as well as that of other researchers, shows better prompting techniques help a lot, but model training is still a huge limitin…

ApplicationsDGX agent

Ethan Mollick discusses research findings showing that while improved prompting techniques provide significant benefits for AI model performance, the underlying model training remains the primary limi

Outlier Smoothing with Closed-Form Rotations for W4A4 Large Language Model Quantization

Model ReleasesDGX agent

arXiv:2511.22316v2 Announce Type: replace Abstract: Large Language Models (LLMs) quantization facilitates deploying LLMs in resource-limited settings, but existing methods that combine incompatible gr

PathPainter: Transferring the Generalization Ability of Image Generation Models to Embodied Navigation

Model ReleasesDGX agent

arXiv:2605.07496v1 Announce Type: new Abstract: Bird's-eye-view (BEV) images have been widely demonstrated to provide valuable prior information for navigation. Given the global information provided b

PolarVLM: Bridging the Semantic-Physical Gap in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.07574v1 Announce Type: new Abstract: Mainstream vision-language models (VLMs) fundamentally struggle with severe optical ambiguities, such as reflections and transparent objects, due to the

ProteinJEPA: Latent prediction complements protein language models

ResearchDGX agent

arXiv:2605.07554v1 Announce Type: cross Abstract: Protein language models are trained primarily with masked language modeling (MLM), which predicts amino-acid identities at masked positions. We ask wh

Reflections and New Directions for Human-Centered Large Language Models

SafetyDGX agent

arXiv:2605.06901v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly shaping the private and professional lives of users, with numerous applications in business, education, fi

Stabilized neural Hamilton--Jacobi--Bellman solvers: Error analysis and applications in model-based reinforcement learning

SafetyDGX agent

arXiv:2605.07116v1 Announce Type: cross Abstract: Physics-informed neural solvers offer a promising route to model-based reinforcement learning in continuous time, where optimal feedback synthesis is

Structured Prototype-Guided Adaptation for EEG Foundation Models

Model ReleasesDGX agent

arXiv:2602.17251v2 Announce Type: replace Abstract: Electroencephalography (EEG) foundation models (EFMs) have shown strong potential for transferable representation learning, yet their adaptation in

TajPersLexon: A Tajik-Persian Lexical Resource and Hybrid Model for Cross-Script Low-Resource NLP

Model ReleasesDGX agent

arXiv:2605.06886v1 Announce Type: new Abstract: This work introduces TajPersLexon, a curated Tajik--Persian parallel lexical resource of 40,112 word and short-phrase pairs for cross-script lexical ret

ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models

Model ReleasesDGX agent

arXiv:2510.01290v2 Announce Type: replace Abstract: The long-output context generation of large reasoning models enables extended chain of thought (CoT) but also drives rapid growth of the key-value (

10 May 2026

what if we name the next model 'goblin' almost worth it to make you all happy...

IndustryDGX agent

Sam Altman jokingly suggested naming OpenAI's next model 'Goblin' as a humorous response to community requests or preferences about model naming. The post appears to be a lighthearted tweet indicating

9 May 2026

The great thing is that the names are so baffling that the most important models OpenAI released were names davinci-002, GPT-3.5, GPT-4, o1-…

Model ReleasesDGX agent

The great thing is that the names are so baffling that the most important models OpenAI released were names davinci-002, GPT-3.5, GPT-4, o1-preview, o3, GPT-5 Pro, and you would never know the ways th

8 May 2026

GPT-5.5 is now available in ComfyUI. OpenAI's latest frontier model — built for reasoning, structured output, and reliable single-pass resul…

Model ReleasesDGX agent

GPT-5.5 is now available in ComfyUI. OpenAI's latest frontier model — built for reasoning, structured output, and reliable single-pass results. → Prompt synthesis → Logic-heavy transformations → Clean

Training models involves many technical and social processes, so prevention of CoT grading has to be built into the process. We’re improving…

Model ReleasesDGX agent

Training models involves many technical and social processes, so prevention of CoT grading has to be built into the process. We’re improving real-time CoT-grading detection, safeguards against acciden

7 May 2026

A Fast Model Counting Algorithm for Two-Variable Logic with Counting and Modulo Counting Quantifiers

ResearchDGX agent

arXiv:2605.03391v1 Announce Type: cross Abstract: Weighted first-order model counting (WFOMC) is a central task in lifted probabilistic inference: It asks for the weighted sum of all models of a first

← Previous
1…114115116117118…1009
Next →