AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,552 results
Safety

Reconsidering Positional Supervision in Masked Diffusion Language Model Training

DGX agent

arXiv:2601.22947v2 Announce Type: replace Abstract: Masked diffusion language models (MDLMs) generate text by unmasking tokens in parallel and have recently emerged as alternatives to autoregressive l

safetyarxiv-cs-cl
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Revisiting Parameter-Based Knowledge Editing in Large Language Models: Theoretical Limits and Empirical Evidence

DGX agent

arXiv:2606.00570v1 Announce Type: cross Abstract: Parameter-based knowledge editing updates the internal knowledge of large language models (LLMs) via localized weight modifications and has attracted

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Self-Regulating Annealing in Heavy-Tailed Diffusion Models

DGX agent

arXiv:2606.01645v1 Announce Type: cross Abstract: Diffusion models have emerged as a leading framework for deep generative modeling. While the standard Gaussian formulation is theoretically convenient

researcharxiv-cs-lg
2 Jun 2026
Model Releases

SentGuard: Sentence-Level Streaming Guardrails for Large Language Models

DGX agent

arXiv:2606.02041v1 Announce Type: new Abstract: Large language models increasingly stream long, reasoning-intensive responses in real time, making when to moderate as critical as whether to moderate.

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

Silent Failures in Federated Personalization of Foundation Models

DGX agent

arXiv:2606.00947v1 Announce Type: cross Abstract: Foundation models are increasingly personalized on decentralized private data through federated learning and are now deployed at scale under growing r

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning

DGX agent

arXiv:2603.08000v2 Announce Type: replace Abstract: Large reasoning models (LRMs) like OpenAI o1 and DeepSeek-R1 achieve high accuracy on complex tasks by adopting long chain-of-thought (CoT) reasonin

model-releasesarxiv-cs-cl
2 Jun 2026
Industry

Step-3.7-Flash from @StepFun_ai is a silent winner. Super impressive results, the best model under 500B params on HF leaderboards.. All whil…

DGX agent

Step-3.7-Flash is a compact language model from StepFun AI that reportedly achieves top-tier performance among models under 500 billion parameters on Hugging Face leaderboards, despite receiving limit

industryclem-delangue--x
2 Jun 2026
Safety

SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models

DGX agent

arXiv:2601.22276v2 Announce Type: replace-cross Abstract: As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributor

safetyarxiv-cs-cv
2 Jun 2026
Model Releases

The Shape of Wisdom: Decision Trajectories in Language Models

DGX agent

arXiv:2606.01202v1 Announce Type: new Abstract: Language models do not simply choose an answer at the output layer. In a 9,000-trajectory MMLU study across Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct,

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models

DGX agent

arXiv:2606.00919v1 Announce Type: new Abstract: Large language models (LLMs) have seen widespread adoption across various domains, yet their reliability is frequently undermined by hallucinations - re

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Trump signs executive order to review AI models before they’re released

DGX agent

President Donald Trump signed an executive order Tuesday creating a 'voluntary framework' for AI companies to share their frontier models with the federal government before they're released 'to promot

model-releasesthe-verge-ai
2 Jun 2026
Research

Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs

DGX agent

arXiv:2506.14003v5 Announce Type: replace Abstract: Machine unlearning (MU) for large language models (LLMs), commonly referred to as LLM unlearning, seeks to remove specific undesirable data or knowl

researcharxiv-cs-lg
2 Jun 2026
Model Releases

Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset

DGX agent

arXiv:2606.02273v1 Announce Type: new Abstract: Understanding subtle driver actions is essential for building reliable driver monitoring systems. Existing visionlanguage models (VLMs) are trained on g

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

We're sponsoring a hackathon to scale down. Hosted by our friends @huggingface and @Gradio, we want working with models to feel like yours a…

DGX agent

We're sponsoring a hackathon to scale down. Hosted by our friends @huggingface and @Gradio, we want working with models to feel like yours again. Small enough that it's inexpensive to run, big enough

model-releasescohere--x
2 Jun 2026
Safety

World Models: A Comprehensive Survey of Architectures, Methodologies, Reasoning Paradigms, and Applications

DGX agent

arXiv:2606.00133v1 Announce Type: new Abstract: World models, internal simulators that learn the structure and dynamics of an environment, have emerged as a central paradigm in the pursuit of artifici

safetyarxiv-cs-lg
2 Jun 2026
Research

3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models

DGX agent

arXiv:2603.07751v2 Announce Type: replace-cross Abstract: Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks

researcharxiv-cs-cl
1 Jun 2026
Hardware

Accelerate LLM model loading and increase context windows with GPUDirect on Amazon FSx for Lustre and TurboQuant

DGX agent

If you’re iterating on deploying large language models (LLMs) on AWS GPU instances, you’ve probably noticed the larger the model to be loaded into GPU High Bandwidth Memory (HBM), the longer the painf

hardwareaws-ml-blog
1 Jun 2026
Model Releases

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education

DGX agent

arXiv:2605.31212v1 Announce Type: cross Abstract: AI systems are increasingly used to support educational content creation, yet it remains unclear whether they can generate outputs that faithfully rep

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Bernini released. Unified Video generation and editing model. Built on Wan-2.2

DGX agent

Bernini is a unified framework for video editing and video generation , built using Wan2.2-A14B as its renderer . The model covers complementary task families that demonstrate its capabilities as a un

model-releasesr-stablediffusion
1 Jun 2026
Model Releases

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

DGX agent

arXiv:2512.19673v3 Announce Type: replace-cross Abstract: Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a unified policy, overlooking their internal mechanisms.

model-releasesarxiv-cs-ai
1 Jun 2026
Research

dgMARK: Decoding-Guided Watermarking for Diffusion Language Models

DGX agent

arXiv:2601.22985v2 Announce Type: replace Abstract: We propose dgMARK, a decoding-guided watermarking method for discrete diffusion language models (dLLMs). Unlike autoregressive models, dLLMs can gen

researcharxiv-cs-lg
1 Jun 2026
Safety

Differentially Private Preference Data Synthesis for Large Language Model Alignment

DGX agent

arXiv:2605.30808v1 Announce Type: cross Abstract: Preference alignment is a crucial post-training step for large language models (LLMs) to ensure their outputs align with human values. However, post-t

safetyarxiv-cs-ai
1 Jun 2026
Research

Domain Adaptation and Reasoning Frameworks in Language Models: A Controlled Experiment with Historical Cosmology

DGX agent

arXiv:2605.30415v1 Announce Type: cross Abstract: We investigate how domain adaptation reshapes explanatory behavior in language models using historical cosmology as a controlled setting. In Phase 1,

researcharxiv-cs-ai
1 Jun 2026
Safety

Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models

DGX agent

arXiv:2509.24319v4 Announce Type: replace-cross Abstract: Large language models can express values in two main ways: (1) intrinsic expression, reflecting the model's inherent values learned during tra

safetyarxiv-cs-ai
1 Jun 2026
Hardware

Elastic ViTs from Pretrained Models without Retraining

DGX agent

arXiv:2510.17700v2 Announce Type: replace Abstract: Vision foundation models achieve remarkable performance but are only available in a limited set of pre-determined sizes, forcing sub-optimal deploym

hardwarearxiv-cs-cv
1 Jun 2026
Applications

Enhancing Computer Vision Model Generalization in Warehouse Facilities: A Case Study on Anomaly Detection in Vertical Material Handling Systems

DGX agent

arXiv:2605.31487v1 Announce Type: new Abstract: Deploying computer vision models in Warehouse Facilities traditionally requires extensive resources for camera mounting, image collection, annotation, t

applicationsarxiv-cs-cv
1 Jun 2026
Research

Fine-Tuning Improves Information Conveyance in Language Models

DGX agent

arXiv:2605.30844v1 Announce Type: cross Abstract: Fine-tuning is often believed to reduce uncertainty and diversity in large language models, but existing analyses overlook output length, a key confou

researcharxiv-cs-ai
1 Jun 2026
Local Ai

Graph Energy Matching: Transport-Aligned Energy-Based Modeling for Graph Generation

DGX agent

arXiv:2603.23398v2 Announce Type: replace-cross Abstract: Generative modeling of discrete data, such as graphs, underpins many scientific and industrial applications, including molecular discovery and

local-aiarxiv-cs-ai
1 Jun 2026
Model Releases

In May, we integrated 11 new models spanning image, 3D, audio, video, and multimodal. The highlights: → Krea 2 — style-first image generatio…

DGX agent

In May, we integrated 11 new models spanning image, 3D, audio, video, and multimodal. The highlights: → Krea 2 — style-first image generation, live as a Partner Node on day one. Competes on how the fr

model-releasescomfyui--x
1 Jun 2026
Research

Language Models Can Resolve Reference Compositionally, But It's Not Their Native Strength: The Case of the Personal Relation Task

DGX agent

arXiv:2605.31480v1 Announce Type: new Abstract: Do neural models, such as Large Language Models, genuinely acquire compositional abilities for interpretation of natural language? When we talk about se

researcharxiv-cs-cl
1 Jun 2026
Model Releases

Many research labs only consider inference efficiency after the fact. Step 3.7 Flash is a 196B MoE model, and built for inference from the s…

DGX agent

Many research labs only consider inference efficiency after the fact. Step 3.7 Flash is a 196B MoE model, and built for inference from the start by @StepFun_ai. Multi-Matrix Factorization Attention (M

model-releasesfireworks-ai--x
1 Jun 2026
Research

Model Monotonicity in Autobidding Auctions: When Do Better Predictions Lead to Better Outcomes?

DGX agent

arXiv:2605.31036v1 Announce Type: cross Abstract: Online advertising platforms rely on machine learning models to predict click-through rates (pCTR) and conversion rates (pCVR) for auction mechanisms.

researcharxiv-cs-lg
1 Jun 2026
Model Releases

MultiPriv: Benchmarking Individual-Level Privacy Reasoning in Vision-Language Models

DGX agent

arXiv:2511.16940v3 Announce Type: replace Abstract: Modern Vision-Language Models (VLMs) pose significant individual-level privacy risks by linking fragmented multimodal data to identifiable individua

model-releasesarxiv-cs-cv
1 Jun 2026
Model Releases

OpenAI frontier models and Codex are now generally available on AWS, giving enterprises a new way to build on Amazon Bedrock with OpenAI thr…

DGX agent

OpenAI frontier models and Codex are now generally available on AWS, giving enterprises a new way to build on Amazon Bedrock with OpenAI through the security, compliance, and governance workflows they

model-releasesopenai--x
1 Jun 2026
Applications

Performance and Complexity Trade-off Optimization of Speech Models During Training

DGX agent

arXiv:2601.13704v3 Announce Type: replace-cross Abstract: In speech machine learning, neural network models are typically designed by choosing an architecture with fixed layer sizes and structure. The

applicationsarxiv-cs-ai
1 Jun 2026
Model Releases

SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders

DGX agent

arXiv:2509.21379v3 Announce Type: replace-cross Abstract: Concept unlearning in diffusion models is hampered by feature splitting, where concepts are distributed across many latent features, making th

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

Sequential Least-Squares Estimators with Fast Randomized Sketching for Linear Statistical Models

DGX agent

arXiv:2509.06856v2 Announce Type: replace-cross Abstract: We propose a novel randomized framework for the estimation problem of large-scale linear statistical models, namely Sequential Least-Squares E

model-releasesarxiv-cs-lg
1 Jun 2026
Safety

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO

DGX agent

arXiv:2605.30789v1 Announce Type: cross Abstract: We identify a new dimension for enhancing rollout diversity in Group Relative Policy Optimization (GRPO) for LLMs. While GRPO relies on diverse rollou

safetyarxiv-cs-ai
1 Jun 2026
Model Releases

Structured interactions improve distributed coordination beyond model scaling in a real-world multi-robot system

DGX agent

arXiv:2605.30383v1 Announce Type: cross Abstract: Scaling individual robot capabilities is common but costly. Here we investigate a system-level design question in real-world multi-robot coordination:

model-releasesarxiv-cs-ai
1 Jun 2026
Model Releases

The Illusion of Generalization in Tabular Language Models

DGX agent

arXiv:2602.04031v2 Announce Type: replace Abstract: Tabular Language Models (TLMs) have been claimed to achieve strong generalization for tabular prediction. We conduct a systematic re-evaluation of T

model-releasesarxiv-cs-lg
1 Jun 2026
Research

Unlearning in Diffusion Models: A Unified Framework with KL Divergence and Likelihood Constraints

DGX agent

arXiv:2605.30825v1 Announce Type: cross Abstract: Unlearning in diffusion models aims to remove undesirable data or concepts while preserving the utility of pretrained models -- two fundamentally conf

researcharxiv-cs-ai
1 Jun 2026
Safety

Vision-Language Models Suppress Female Representations Under Ambiguous Input

DGX agent

arXiv:2605.31556v1 Announce Type: cross Abstract: Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far l

safetyarxiv-cs-ai
1 Jun 2026
Hardware

Welcome NVIDIA Cosmos 3: The First Open Omni-model for Physical AI Reasoning and Action

DGX agent

NVIDIA Cosmos 3 is an open-source omni-model designed for physical AI reasoning and action tasks, representing an advancement in multimodal AI systems. The model integrates multiple modalities to enab

hardwarehugging-face
1 Jun 2026
Safety

World Action Verifier: Self-Improving World Models via Forward-Inverse Asymmetry

DGX agent

arXiv:2604.01985v2 Announce Type: replace-cross Abstract: General-purpose world models promise scalable policy evaluation, optimization, and planning, yet achieving the required level of robustness re

safetyarxiv-cs-ai
1 Jun 2026
Local Ai

Stable Diffusion model recommendations for faster and cleaner outputs in 2026?

DGX agent

A Reddit discussion seeking Stable Diffusion model recommendations for faster and cleaner image outputs, addressing the reality that no single best model exists as the right choice depends on hardware

local-air-stablediffusion
31 May 2026
Model Releases

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, yo…

DGX agent

We need more coding and agent traces public sharing to build datasets and better open source models! Lots of people contributing already, you should share yours too! https://huggingface.co/datasets?se

model-releasesclem-delangue--x
31 May 2026
Safety

is having a four month lead a sustainable multitrillion dollar business model?

DGX agent

is having a four month lead a sustainable multitrillion dollar business model? We took another look at the capability gap between open-weight and proprietary models. Since the start of the year, open-

safetygary-marcus--x
30 May 2026
Agents

A Deep Learning Model of Mental Rotation Informed by Interactive VR Experiments

DGX agent

arXiv:2512.13517v2 Announce Type: replace-cross Abstract: Mental rotation -- the ability to compare objects seen from different viewpoints -- is a fundamental example of mental simulation and spatial

agentsarxiv-cs-lg
29 May 2026
← Previous
1…136137138139140…1262
Next →