AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
60,588 results
28 May 2026

Segment to Focus: Guiding Latent Action Models in the Presence of Distractors

AgentsDGX agent

arXiv:2602.02259v2 Announce Type: replace-cross Abstract: Latent action models (LAMs) offer a promising path to pre-training embodied agents on large amounts of action-free video. They infer latent ac

Surrogate modeling for Bayesian optimization beyond a single Gaussian process

ApplicationsDGX agent

arXiv:2205.14090v2 Announce Type: replace-cross Abstract: Bayesian optimization (BO) has well-documented merits for optimizing black-box functions with an expensive evaluation cost. Such functions eme

Tell Me a Story! Narrative-Driven XAI with Large Language Models

ResearchDGX agent

arXiv:2309.17057v3 Announce Type: replace Abstract: In many AI applications today, the predominance of black-box machine learning models, due to their typically higher accuracy, amplifies the need for

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Text-Only Data Synthesis for Vision Language Model Training

ResearchDGX agent

arXiv:2503.22655v2 Announce Type: replace Abstract: Training vision-language models (VLMs) typically requires large-scale, high-quality image-text pairs, but collecting or synthesizing such data is co

Today, we're releasing LFM2.5-8B-A1B, a device-optimized model designed to power real-life applications on phones, laptops, PCs, robots, and…

HardwareDGX agent

Today, we're releasing LFM2.5-8B-A1B, a device-optimized model designed to power real-life applications on phones, laptops, PCs, robots, and fast & lightweight server-side use-cases. > 8B MoE, 1.5B ac

Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs

SafetyDGX agent

arXiv:2605.27894v1 Announce Type: new Abstract: Video-Language Models (VLMs) have demonstrated impressive multi-modal reasoning capabilities across diverse computer vision applications. However, these

Training Azerbaijani language models on Amazon SageMaker AI

ApplicationsDGX agent

Azercell Telecom LLC, Azerbaijan's leading telecommunications provider, wanted to build an Azerbaijani large language model (LLM) on Amazon SageMaker AI for telecom use cases and a customer-facing cha

Training Stratigraphy: Persistent Behavioral Artifacts in Large Language Models Observed Through Longitudinal AI-Human Interaction

SafetyDGX agent

arXiv:2605.28102v1 Announce Type: new Abstract: Large language models trained with Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI exhibit persistent behavioral patterns that s

Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models

SafetyDGX agent

arXiv:2605.27376v1 Announce Type: cross Abstract: While prompt-based text-to-speech (TTS) models enable natural language-driven speaking style control, they often provide limited fine-grained control

When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models

ResearchDGX agent

arXiv:2605.28181v1 Announce Type: new Abstract: Diffusion language models decode text by iteratively denoising masked token sequences, making the choice of which positions to decode a central inferenc

When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs

ResearchDGX agent

arXiv:2605.28346v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly evaluated for whether they identify the right visual content, but little is known about whether they expr

When pre-training hurts LoRA fine-tuning: a dynamical analysis via single-index models

SafetyDGX agent

arXiv:2602.02855v2 Announce Type: replace Abstract: Pre-training on a source task is usually expected to facilitate fine-tuning on similar downstream problems. In this work, we mathematically show tha

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

TutorialsDGX agent

arXiv:2605.28132v1 Announce Type: new Abstract: Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this,

Would you like to join the research effort on JEPA and World Models easily? After a full year of hard work, we’re excited to finally release…

ResearchDGX agent

Would you like to join the research effort on JEPA and World Models easily? After a full year of hard work, we’re excited to finally release stable-worldmodel: an open-source, scalable platform built

27 May 2026

Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models

SafetyDGX agent

arXiv:2605.26491v1 Announce Type: cross Abstract: Preference optimization has emerged as an efficient alternative to online reinforcement learning from human feedback (RLHF) for aligning text-to-image

Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models

ResearchDGX agent

arXiv:2605.27311v1 Announce Type: new Abstract: Chart question-answering (QA) benchmarks aim to pose questions that require visual reasoning to correctly answer, but models can often reach solutions t

CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning

Model ReleasesDGX agent

arXiv:2605.26967v1 Announce Type: new Abstract: Existing video captioning methods struggle to balance visual fidelity and redundancy: holistic captions are compact but lose fine-grained evidence, wher

ContextGuard: Structured Self-Auditing for Context Learning in Language Models

ResearchDGX agent

arXiv:2605.26827v1 Announce Type: cross Abstract: Recent benchmarks reveal that despite strong reasoning capabilities, large language models (LLMs) still struggle to faithfully apply complex contextua

Cultural Value Alignment Via Latent Activation Steering in Large Language Models

SafetyDGX agent

arXiv:2605.26365v1 Announce Type: new Abstract: Large Language Models (LLMs) often exhibit homogenized cultural perspectives. While the World Values Survey (WVS) provides a gold standard for mapping h

Edge AI Deployment Beyond Models: A BSP-Aware Systems Framework for Industrial Embedded Platforms

Local AiDGX agent

arXiv:2605.26119v1 Announce Type: cross Abstract: Industrial Edge AI programs often begin with the model and only later confront the platform. That sequencing is attractive because it allows early dem

Emergent Causal-Geometric Dynamics Across Depth in Large Language Models

ResearchDGX agent

arXiv:2602.04931v2 Announce Type: replace-cross Abstract: Geometric analyses of large language model (LLM) representations reveal structured variation across depth but remain fundamentally correlation

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights

SafetyDGX agent

arXiv:2501.06708v5 Announce Type: replace-cross Abstract: Large-scale web-crawled datasets contain noise, bias, and irrelevant information, necessitating data selection techniques. Existing methods de

Google has the only true Omni model, but the elements aren't hooked up. It appears it can take in & output audio, images. video, songs, text…

ApplicationsDGX agent

Google has the only true Omni model, but the elements aren't hooked up. It appears it can take in & output audio, images. video, songs, text, code, etc. But right now each type of output is separate.

Great to see @poolsideai (US lab) committing to open sourcing their foundation models going forward Laguna is an interesting release, check …

IndustryDGX agent

Great to see @poolsideai (US lab) committing to open sourcing their foundation models going forward Laguna is an interesting release, check it out @Shaughnessy119 https://poolside.ai/blog/introducing-

If the Founder of Hugging Face asks, you gotta do it. Models and dataset now live: https://huggingface.co/papers/2605.22391 Also built an ex…

IndustryDGX agent

If the Founder of Hugging Face asks, you gotta do it. Models and dataset now live: https://huggingface.co/papers/2605.22391 Also built an explorer: https://huggingface.co/spaces/Kaikaku/epicure-explor

Learning Energy-Based Models from Stochastic Interpolants using Spatiotemporal Differences

TutorialsDGX agent

arXiv:2605.26850v1 Announce Type: new Abstract: Learning an energy-based model from data samples is a central problem in machine learning. Many recent and popular methods, such as denoising score matc

Learning to Diagnose and Correct Errors: Towards Moral Sensitivity Acquisition in Large Language Models

TutorialsDGX agent

arXiv:2601.03079v4 Announce Type: replace Abstract: Moral sensitivity is the most fundamental capability underlying human moral competence. Although many approaches aim to align large language models

Left-Right Symmetry Breaking in CLIP-style Vision-Language Models Trained on Synthetic Spatial-Relation Data

ResearchDGX agent

arXiv:2601.12809v2 Announce Type: replace-cross Abstract: Spatial understanding remains a key challenge in vision-language models. Yet it is still unclear whether such understanding is truly acquired,

LUCoS: Latent Unsupervised Context Selection for Tabular Foundation Models

ResearchDGX agent

arXiv:2605.27254v1 Announce Type: cross Abstract: Selecting which instances to label is a key challenge in low-label tabular learning. For recent Tabular Foundation Models such as TabPFN, context sele

Model Unlearning Objectives Vary for Distinct Language Functions

TutorialsDGX agent

arXiv:2605.26454v1 Announce Type: new Abstract: Large language models (LLMs) learn undesirable properties during pretraining, including dangerous knowledge and toxic text generation. Just as post-trai

Olaf-World: Orienting Latent Actions for Video World Modeling

SafetyDGX agent

arXiv:2602.10104v2 Announce Type: replace-cross Abstract: Scaling action-controllable world models is limited by the scarcity of action labels. While latent action learning promises to extract control

Respecting Modality Gap in Post-hoc Out-of-distribution Detection with Pre-trained Vision-Language Models

ResearchDGX agent

arXiv:2605.26661v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection has emerged as a popular technique to enhance the reliability of machine learning models by identifying unexpected

SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability

Model ReleasesDGX agent

arXiv:2605.26788v1 Announce Type: cross Abstract: Large language models (LLMs) achieve impressive performance when a task is fully specified in a single turn, yet the same models lose up to 39% of tha

SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning

Model ReleasesDGX agent

arXiv:2603.28730v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have shown impressive capabilities across diverse tasks, motivating efforts to leverage these models to supervis

The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models

Local AiDGX agent

arXiv:2605.26128v1 Announce Type: new Abstract: Production LLM systems increasingly require machine-readable outputs: JSON objects, typed traces, regex-constrained fields, and tool-call schemas. This

To model human linguistic prediction, make LLMs less superhuman

SafetyDGX agent

arXiv:2510.05141v2 Announce Type: replace Abstract: When we read, we make predictions about upcoming words; these predictions influence our reading behavior. The success of large language models (LLMs

Wow. It looks like the @XiaomiMiMo v2.5 model is insanely good value :O (Price for each prompt shown after each answer. Context includes >40…

TutorialsDGX agent

Jeremy Howard comments on the Xiaomi MiMo v2.5 model, highlighting its exceptional value proposition and cost-effectiveness for prompt processing. The post appears to include comparative pricing data

26 May 2026

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

SafetyDGX agent

arXiv:2605.26026v1 Announce Type: cross Abstract: Light sheet fluorescence microscopy (LSM) enables high-resolution, three-dimensional (3D) imaging of biological specimens, providing rich volumetric d

AutoSOTA: An End-to-End Automated Research System for State-of-the-Art AI Model Discovery

AgentsDGX agent

arXiv:2604.05550v2 Announce Type: replace Abstract: Artificial intelligence research increasingly depends on prolonged cycles of reproduction, debugging, and iterative refinement to achieve State-Of-T

Capability and Robustness Cannot Both Be Free: An Information-Theoretic Bound for Vision-Language-Action Models

SafetyDGX agent

arXiv:2605.25889v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are increasingly deployed on real robots, where each predicted action is executed and each failure carries a safet

Comfy will be at @aionthelot Our CEO @yoland_yan is joining a panel of AI video leaders for a progress report on the state of video models -…

Local AiDGX agent

Comfy will be at @aionthelot Our CEO @yoland_yan is joining a panel of AI video leaders for a progress report on the state of video models - where they excel, where they fall short, and what's coming

Confidence Calibration in Large Language Models

ResearchDGX agent

arXiv:2605.23909v1 Announce Type: new Abstract: We investigate the calibration of large language models' (LLMs') confidence across diverse tasks. The results of our preregistered study show that the c

Dynamic Optimization and Safety Indicator Injection for Jailbreaking Text-to-Image Models with Multimodal Safety Filters

SafetyDGX agent

arXiv:2505.18979v2 Announce Type: replace Abstract: Text-to-image (T2I) models can generate not-safe-for-work (NSFW) content, motivating multi-stage safety pipelines with both text and image filters.

E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving

SafetyDGX agent

arXiv:2512.04733v2 Announce Type: replace-cross Abstract: End-to-end autonomous driving (AD) systems increasingly adopt vision-language-action (VLA) models, yet they typically ignore the passenger's e

ERNIE-Image Technical Report

Model ReleasesDGX agent

arXiv:2605.25347v1 Announce Type: cross Abstract: We introduce ERNIE-Image, an open-source text-to-image generation model built upon an 8B single-stream DiT architecture. ERNIE-Image aims to bridge th

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

TutorialsDGX agent

arXiv:2605.25477v1 Announce Type: cross Abstract: The ability to efficiently and reliably learn new tasks has been a foundational challenge in robotics. Vision-Language-Action (VLA) models have demons

Hadamard Representation: Scaffolding Performance Across Model-free RL

Model ReleasesDGX agent

arXiv:2406.09079v5 Announce Type: replace Abstract: Deep reinforcement learning agents progressively lose representational capacity during training: neurons become dormant, removing active capacity fr

Hypothesis Generation and Inductive Inference in Children and Language Models

ApplicationsDGX agent

arXiv:2605.24528v1 Announce Type: new Abstract: Real world decision-making requires constructing mental models under uncertainty over evidence, over the underlying causal rules, and over the state of

Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models

SafetyDGX agent

arXiv:2605.24550v1 Announce Type: new Abstract: Fine-tuning-as-a-Service (FaaS) enables personalization of large language models (LLMs), but it can weaken safety-alignment under harmful fine-tuning at

// Language Models Need Sleep // Let your agents 'sleep', folks. On a serious note, this is a fascinating paper on getting the most from lon…

TutorialsDGX agent

// Language Models Need Sleep // Let your agents 'sleep', folks. On a serious note, this is a fascinating paper on getting the most from long-horizon agents. Here is the problem with agents today: Att

Learning to Reason Efficiently with A* Post-Training

Model ReleasesDGX agent

arXiv:2605.24597v1 Announce Type: new Abstract: Many applications of large language models (LLMs) require deductive reasoning, yet models frequently produce incorrect or redundant inference steps. We

Logic-Guided Vector Fields for Constrained Generative Modeling

ResearchDGX agent

arXiv:2602.02009v2 Announce Type: replace Abstract: Neuro-symbolic systems aim to combine the expressive structure of symbolic logic with the flexibility of neural learning; yet, generative models typ

MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models

SafetyDGX agent

arXiv:2605.26004v1 Announce Type: cross Abstract: Instruction tuning of large vision-language models (LVLMs) increasingly depends on massive multimodal corpora, yet these datasets contain samples with

MCPXKIT: The Unified Toolkit for Analyzing Model Context Protocol Security

AgentsDGX agent

arXiv:2508.12538v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) has emerged as a universal standard that enables AI agents to seamlessly connect with external tools, signifi

Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration

ResearchDGX agent

arXiv:2605.24957v1 Announce Type: new Abstract: The generation of factually incorrect objects, commonly known as object hallucination, remains a persistent challenge in Large Vision-Language Models (L

OpenRouter raised 113M led by CapitalG, a source says at a 1.3B valuation, and now processes 25T tokens across 400+ models weekly, up from 5T six months ago (Michael J. de la Merced/New York Times)

IndustryDGX agent

Michael J. de la Merced / New York Times: OpenRouter raised 113M led by CapitalG, a source says at a 1.3B valuation, and now processes 25T tokens across 400+ models weekly, up from 5T six months ago —

Regional Condition Custom Node for Anima model

Local AiDGX agent

Anima is a 2 billion parameter text-to-image model focused mainly on anime concepts and styles, but also capable of generating other non-photorealistic content. A Regional Condition Custom Node for An

Side-by-side Comparison Amplifies Dialect Bias in Language Models

SafetyDGX agent

arXiv:2605.24384v1 Announce Type: cross Abstract: Language models (LMs) can exhibit systematic biases against speakers based on variations in their dialects, even in the absence of a dialect label, a

Sources: Fireworks AI, which helps companies run AI models, is in talks to raise funding at a 15B valuation after being valued at 4B in October 2025 (Rebecca Torrence/Bloomberg)

IndustryDGX agent

Rebecca Torrence / Bloomberg: Sources: Fireworks AI, which helps companies run AI models, is in talks to raise funding at a 15B valuation after being valued at 4B in October 2025 — Fireworks AI, a sta

Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation

ResearchDGX agent

arXiv:2605.26111v1 Announce Type: cross Abstract: Subject-driven image generation aims to synthesize new images that preserve the identity of the given subject while following textual instructions. Ex

← Previous
1…182183184185186…1010
Next →