AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
59,929 results
Model Releases

MACD: Model-Aware Contrastive Decoding via Counterfactual Data

DGX agent

arXiv:2602.01740v3 Announce Type: replace Abstract: Video language models (Video-LLMs) are prone to hallucinations, generating plausible but ungrounded content when visual evidence is weak, ambiguous,

model-releasesarxiv-cs-ai
8 Jun 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

DGX agent

arXiv:2606.06696v1 Announce Type: cross Abstract: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models

DGX agent

arXiv:2606.06853v1 Announce Type: cross Abstract: The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLM

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

The Geometry of Representational Failures in Vision Language Models

DGX agent

arXiv:2602.07025v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) exhibit puzzling failures in multi-object visual tasks, such as hallucinating non-existent elements or failing t

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

TokaMind: A Multi-Modal Transformer Foundation Model for Tokamak Plasma Dynamics

DGX agent

arXiv:2602.15084v2 Announce Type: replace-cross Abstract: We present TokaMind, to our knowledge the first open-source foundation model for tokamak plasma dynamics, based on a Multi-Modal Transformer (

model-releasesarxiv-cs-ai
8 Jun 2026
Model Releases

Learning What Matters: Probabilistic Task Selection via Mutual Information for Model Finetuning

DGX agent

arXiv:2507.12612v3 Announce Type: replace-cross Abstract: Supervised fine-tuning performance for large language models depends strongly on how training budget is distributed across a heterogeneous set

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

LLMCodec: Adapting Video Codecs for Efficient Weight Compression of Large Language Models

DGX agent

arXiv:2606.05861v1 Announce Type: cross Abstract: The rapid development of large language models(LLMs) has led to remarkable advances in natural language processing. However, the increasing scale of t

model-releasesarxiv-cs-ai
6 Jun 2026
Model Releases

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models

DGX agent

arXiv:2606.05434v1 Announce Type: cross Abstract: Group Relative Policy Optimisation (GRPO) has emerged as an effective reinforcement-learning algorithm for aligning language models on reasoning tasks

model-releasesarxiv-cs-ai
6 Jun 2026
Research

Severity-Aware Curriculum Learning with Multi-Model Response Selection for Medical Text Generation

DGX agent

arXiv:2606.05510v1 Announce Type: new Abstract: Telehealth systems have become increasingly important for delivering accessible and timely medical information. Existing large language models often str

researcharxiv-cs-ai
6 Jun 2026
Model Releases

Evaluating Stochastic Collapse and Implicit Bias in Multimodal Large Language Models

DGX agent

arXiv:2606.05874v1 Announce Type: new Abstract: Current evaluations for Multimodal Large Language Models (MLLMs) overwhelmingly focus on utility-driven objectives, leaving model behavior under logic-n

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Faithful, Enriched, and Precise: Benchmarking Natural-Science Illustration Generation by T2I models

DGX agent

arXiv:2606.05949v1 Announce Type: new Abstract: Scientific illustrations are essential tools for communicating research findings, especially in natural science, where they visualize complex concepts a

model-releasesarxiv-cs-cv
5 Jun 2026
Model Releases

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

DGX agent

arXiv:2606.05177v1 Announce Type: new Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and

model-releasesarxiv-cs-cl
5 Jun 2026
Research

Measuring the sensitivity of LLM-based structured extraction to prompt, model, and schema choices in clinical discharge summaries

DGX agent

arXiv:2606.05970v1 Announce Type: new Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream con

researcharxiv-cs-cl
5 Jun 2026
Model Releases

Nemotron 3 Ultra is now available for Pro and Max subscribers on Perplexity and Computer. It's @nvidia's new open model built for long-runni…

DGX agent

Nvidia's Nemotron 3 Ultra, an open-source model designed for long-context tasks, is now available to Pro and Max subscribers on Perplexity and Computer. This new model represents Nvidia's latest contr

model-releasesperplexity--x
5 Jun 2026
Model Releases

The Granularity Gap: A Multi-Dimensional Longitudinal Audit of Sycophancy in Gemini Models

DGX agent

arXiv:2606.05183v1 Announce Type: new Abstract: Large language models are increasingly deployed as high-stakes advisors, yet standard alignment benchmarks treat sycophancy as a binary failure mode. We

model-releasesarxiv-cs-cl
5 Jun 2026
Model Releases

Using Large Language Models to Support High Volume Application Review for an Undergraduate Research Program

DGX agent

arXiv:2606.05564v1 Announce Type: new Abstract: Undergraduate research programs such as the Summer Undergraduate Research Fellowship (SURF) at Purdue University receive thousands of applications every

model-releasesarxiv-cs-cl
5 Jun 2026
Safety

Value-and-Structure Alignment for Routing-Consistent Quantization of Mixture-of-Experts Models

DGX agent

arXiv:2606.05688v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale foundation models efficiently by activating only a subset of experts for each token, but their large number of exp

safetyarxiv-cs-cl
5 Jun 2026
Research

What Makes Two Language Models Think Alike?

DGX agent

arXiv:2406.12620v3 Announce Type: replace Abstract: Do architectural and training differences influence the way models represent and process language? Traditional similarity metrics tell us whether tw

researcharxiv-cs-cl
5 Jun 2026
Safety

A Unified Geometric Space for Topological Alignment Between Transformer-Based Models and Human Brain Networks

DGX agent

arXiv:2510.24342v2 Announce Type: replace Abstract: Prior brain-AI alignment studies are typically constrained by specific inputs and tasks, limiting their ability to capture organizational properties

safetyarxiv-cs-ai
4 Jun 2026
Model Releases

Can Large Language Models Generalize Procedures Across Representations?

DGX agent

arXiv:2602.03542v2 Announce Type: replace Abstract: Large language models (LLMs) are trained and tested extensively on symbolic representations such as code and graphs, yet real-world user tasks are o

model-releasesarxiv-cs-cl
4 Jun 2026
Model Releases

LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion

DGX agent

arXiv:2602.12215v2 Announce Type: replace Abstract: Recent robot foundation models largely rely on large-scale behavior cloning, which imitates expert actions but discards transferable dynamics knowle

model-releasesarxiv-cs-ro
4 Jun 2026
Model Releases

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training

DGX agent

arXiv:2506.05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention. Although widely adopted, transfo

model-releasesarxiv-cs-ai
4 Jun 2026
Model Releases

Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from k-Parity

DGX agent

arXiv:2601.22450v2 Announce Type: replace-cross Abstract: Masked Diffusion Language Models have recently emerged as a powerful generative paradigm, yet their generalization properties remain understud

model-releasesarxiv-cs-ai
4 Jun 2026
Safety

Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Model Enhancement

DGX agent

arXiv:2412.01282v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) bring powerful understanding and reasoning capabilities to multimodal tasks. Meanwhile, the great need for capab

safetyarxiv-cs-ai
3 Jun 2026
Model Releases

Can Structural Cues Save LLMs? Evaluating Language Models in Massive Document Streams

DGX agent

arXiv:2603.19250v2 Announce Type: replace Abstract: Evaluating language models in streaming environments is critical, yet underexplored. Existing benchmarks either focus on single complex events or pr

model-releasesarxiv-cs-cl
3 Jun 2026
Local Ai

Coupled Local and Global World Models for Efficient First Order RL

DGX agent

arXiv:2602.06219v2 Announce Type: replace-cross Abstract: World models offer a promising avenue for more faithfully capturing complex dynamics, including contacts and non-rigidity, as well as complex

local-aiarxiv-cs-ai
3 Jun 2026
Model Releases

.@GoogleDeepMind's Gemma 4 - 12B is available on Ollama! Chat: ollama run gemma4:12b-mlx Hermes Agent: ollama launch hermes --model gemma4:1…

DGX agent

.@GoogleDeepMind's Gemma 4 - 12B is available on Ollama! Chat: ollama run gemma4:12b-mlx Hermes Agent: ollama launch hermes --model gemma4:12b-mlx Claude Code: ollama launch claude --model gemma4:12b-

model-releasesollama--x
3 Jun 2026
Tutorials

Grounding Functional Similarity by Invariance-Aware Model Stitching

DGX agent

arXiv:2505.20142v2 Announce Type: replace Abstract: In deep learning, functional similarity evaluation quantifies the extent to which independently trained models learn similar input--output relations

tutorialsarxiv-cs-lg
3 Jun 2026
Model Releases

Investigating Adversarial Robustness of Multi-modal Large Language Models

DGX agent

arXiv:2606.03713v1 Announce Type: new Abstract: Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

QUIVER: Quantum-Informed Views for Enhanced Representations in Large ML Models

DGX agent

arXiv:2606.02785v1 Announce Type: new Abstract: Large machine learning models benefit substantially from multimodal inputs that provide a complementary view of the same example. We introduce QUIVER (Q

model-releasesarxiv-cs-lg
3 Jun 2026
Research

Self-Soupervision: Cooking Model Soups without Labels

DGX agent

arXiv:2602.02890v2 Announce Type: replace Abstract: Model soups are strange and strangely effective combinations of parameters. They take a model (the stock), fine-tune it into multiple models (the in

researcharxiv-cs-lg
3 Jun 2026
Model Releases

Spike-Aware C++ INT8 Inference for Sparse Spiking Language Models on Commodity CPUs

DGX agent

arXiv:2606.03026v1 Announce Type: cross Abstract: Spiking language models expose activation sparsity that dense Transformer runtimes do not directly exploit. This paper studies that property from a sy

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

DGX agent

arXiv:2606.00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or

model-releasesarxiv-cs-ai
2 Jun 2026
Industry

Automatic behind the scene routing in user interfaces (instead of model picker) will redistribute value capture and usage towards many more …

DGX agent

Automatic behind the scene routing in user interfaces (instead of model picker) will redistribute value capture and usage towards many more models than just frontier ones (especially towards open-sour

industryclem-delangue--x
2 Jun 2026
Model Releases

Collaborative Few-Step Distillation and Low-Bit Quantization for Wan2.2 Dual-Expert Video Diffusion Models

DGX agent

arXiv:2606.00658v1 Announce Type: cross Abstract: Large video diffusion models achieve strong visual quality but remain expensive to deploy because each sample requires many denoising steps and a larg

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Computation-Aware Kalman Filtering with Model Selection for Neural Dynamics

DGX agent

arXiv:2606.01468v1 Announce Type: cross Abstract: Due to their explicit priors and ability to model uncertainty, Bayesian methods have played a major role in dynamical latent variable modeling of sing

researcharxiv-cs-ai
2 Jun 2026
Research

EPIC: Efficient and Parallel Inference under CFG Constraints for Diffusion Language Models

DGX agent

arXiv:2606.00722v1 Announce Type: cross Abstract: Controlling language model outputs is essential for ensuring structural validity, reliability, and downstream usability, and diffusion language models

researcharxiv-cs-ai
2 Jun 2026
Model Releases

Evaluating the Reversal Curse in Model Editing

DGX agent

arXiv:2310.10322v3 Announce Type: replace Abstract: Large language models (LLMs) are prone to hallucinate unintended text due to false or outdated knowledge. Since retraining LLMs is resource intensiv

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

From Segments to Scenes: Temporal Understanding in Autonomous Driving via Vision-Language Model

DGX agent

arXiv:2512.05277v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

KliniskVestBERT: BERT Model Specialised to Norwegian Clinical Texts

DGX agent

arXiv:2606.01904v1 Announce Type: cross Abstract: The increasing application of Natural Language Processing (NLP) in healthcare demands language models specifically attuned to the complexities of clin

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Med-URWKV{ag}: Toward Enhanced Pretrained Pure VRWKV Models for Medical Image Segmentation

DGX agent

arXiv:2506.10858v2 Announce Type: replace-cross Abstract: Medical image segmentation is a fundamental task in computer-aided diagnosis and treatment. Existing approaches based on CNNs, ViTs, Mamba, an

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

Microsoft's new MAI models

DGX agent

Microsoft announced two new text LLMs this morning - MAI-Thinking-1 (reasoning, 35B parameters, available to 'select early partners') and MAI-Code-1-Flash (5B parameters, 'purpose-built for GitHub Cop

model-releasessimon-willison
2 Jun 2026
Agents

MIRROR: A Multi-Agent Framework with Iterative Adaptive Revision and Hierarchical Retrieval for Optimization Modeling in Operations Research

DGX agent

arXiv:2602.03318v3 Announce Type: replace Abstract: Operations Research (OR) relies on expert-driven modeling-a slow and fragile process ill-suited to novel scenarios. While large language models (LLM

agentsarxiv-cs-cl
2 Jun 2026
Research

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

DGX agent

arXiv:2602.10623v2 Announce Type: replace-cross Abstract: Reward models learned from human preferences are central to aligning large language models (LLMs) via reinforcement learning from human feedba

researcharxiv-cs-ai
2 Jun 2026
Model Releases

Model-Based Quality Assessment for Massively Multilingual Parallel Data

DGX agent

arXiv:2606.00285v1 Announce Type: new Abstract: Large-scale multilingual bitext often contains two distinct problems: non-parallel sentence pairs and low-quality translations. We decompose model-based

model-releasesarxiv-cs-cl
2 Jun 2026
Safety

On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance

DGX agent

arXiv:2606.00467v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-intern

safetyarxiv-cs-ai
2 Jun 2026
Research

Random Erasing vs. Model Inversion: A Promising Defense or a False Hope?

DGX agent

arXiv:2409.01062v4 Announce Type: replace-cross Abstract: Model Inversion (MI) attacks pose a significant privacy threat by reconstructing private training data from machine learning models. While exi

researcharxiv-cs-cv
2 Jun 2026
Model Releases

Self-Healing Agentic Orchestrators for Reliable Tool-Augmented Large Language Model Systems

DGX agent

arXiv:2606.01416v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents rely on orchestration layers that coordinate planning, retrieval, tool invocation, validation, memory,

model-releasesarxiv-cs-ai
2 Jun 2026
← Previous
1…6263646566…1249
Next →