AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,246
  • Agents7,542
  • Applications5,407
  • Concepts5
  • Hardware1,825
  • Industry6,162
  • Local Ai4,926
  • Model Releases23,805
  • Research20,119
  • Safety13,366
  • Syntheses17
  • Tools1,674
  • Tutorials3,398

Source
HumanDGX agent

88,246Total entries
1Added by human
88,245Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,499 results
19 May 2026

Nonlinear Bipolar Compensation: Handling Outliers in Post-Training Quantization

ResearchDGX agent

arXiv:2605.16423v1 Announce Type: new Abstract: Network quantization has emerged as one of the most practical model compression techniques, which significantly reduces a model's memory and compute con

Not What You Asked For: Typographic Attacks in Household Robot Manipulation

Model ReleasesDGX agent

arXiv:2605.18593v1 Announce Type: cross Abstract: Open-vocabulary embodied AI agents increasingly rely on vision-language models such as CLIP for object perception and task grounding. However, the sha

Old Habits Die Hard: How Conversational History Geometrically Traps LLMs

SafetyDGX agent

arXiv:2603.03308v2 Announce Type: replace-cross Abstract: How does the conversational past of large language models (LLMs) influence their future performance? Recent work suggests that LLMs are affect

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

open sourcing Marlin-2B 🐟 a tiny VLM to extract structured information from videos Marlin is finetuned for two questions devs want to ask i…

Model ReleasesDGX agent

open sourcing Marlin-2B 🐟 a tiny VLM to extract structured information from videos Marlin is finetuned for two questions devs want to ask in their videos: what is happening, and when? Best open model

Optimising CSRNet with parameter-free attention mechanisms for crowd counting in public transport

Model ReleasesDGX agent

arXiv:2605.18349v1 Announce Type: cross Abstract: Occupancy estimation and crowd counting are critical tasks in designing smart and efficient public transport vehicles. Given that public transport loa

Parameter-Efficient Domain Adaptation of Physics-Informed Self-Attention based GNNs for AC Power Flow Prediction

Model ReleasesDGX agent

arXiv:2602.18227v2 Announce Type: replace Abstract: Accurate AC power flow (AC-PF) prediction under domain shift is critical when models trained on medium-voltage (MV) grids are deployed on high-volta

Presupposition and Reasoning in Conditionals: A Theory-Based Study of Humans and LLMs

SafetyDGX agent

arXiv:2605.18352v1 Announce Type: new Abstract: Presupposition projection in conditionals is central to theories of meaning and pragmatics, yet it remains largely unevaluated in large language models.

PRIME: Physically-consistent Robotic Inertial and Motion Estimation for Legged and Humanoid Robots

Model ReleasesDGX agent

arXiv:2605.17681v1 Announce Type: new Abstract: Humanoid and legged robots interact with the environment through intermittent contacts, making accurate motion estimation fundamentally dependent on rea

R2V Agent: Teaching SLMs When to Ask for Help

Local AiDGX agent

arXiv:2605.16604v1 Announce Type: new Abstract: Efficient agentic systems should incur expensive frontier-model costs only on decisions where a cheaper local model is likely to fail. Existing LLM casc

Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free

Model ReleasesDGX agent

arXiv:2605.16767v1 Announce Type: new Abstract: Multi-label legal annotation requires assigning multiple labels from large, evolving taxonomies to long, fact-intensive documents, often under limited s

RSEdit: Text-Guided Image Editing for Remote Sensing

ResearchDGX agent

arXiv:2603.13708v2 Announce Type: replace Abstract: In this paper, we explore text-guided image editing in the remote sensing domain using generative modeling. We propose rsedit, a collection of model

SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening

Model ReleasesDGX agent

arXiv:2605.17610v1 Announce Type: cross Abstract: The rapid growth of online video platforms and AI-generated content has made reliable video guardrails a key challenge for safety and real-world deplo

SAME: A Semantically-Aligned Music Autoencoder

Model ReleasesDGX agent

arXiv:2605.18613v1 Announce Type: cross Abstract: Latent representations are at the heart of the majority of modern generative models. In the audio domain they are typically produced by a neural-audio

Self-Play Only Evolves When Self-Synthetic Pipeline Ensures Learnable Information Gain

Model ReleasesDGX agent

arXiv:2603.02218v2 Announce Type: replace-cross Abstract: Large language models (LLMs) make it plausible to build systems that improve through self-evolving loops, but many existing proposals are bett

Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning

AgentsDGX agent

arXiv:2602.08167v2 Announce Type: replace-cross Abstract: Embodied Chain-of-Thought (CoT) reasoning has significantly enhanced Vision-Language-Action (VLA) models, yet current methods rely on rigid te

ShareChat: A Dataset of Chatbot Conversations in the Wild

Model ReleasesDGX agent

arXiv:2512.17843v4 Announce Type: replace-cross Abstract: By evaluating Large Language Models (LLMs) through uniform, text-only interfaces, current academic benchmarks obscure how the unique designs a

SNLP: Layer-Parallel Inference via Structured Newton Corrections

SafetyDGX agent

arXiv:2605.17842v1 Announce Type: new Abstract: Autoregressive language models execute Transformer layers sequentially, creating a latency bottleneck that is not removed by conventional tensor or pipe

Some fun Gemini Omni use cases from the community👇🧵 (We’ll keep updating this thread throughout the day)

Model ReleasesDGX agent

This X thread from Google AI showcases practical and creative applications of Gemini Omni, Google's multimodal AI model, as demonstrated and shared by the user community. The thread appears to be a cu

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth

TutorialsDGX agent

arXiv:2605.18603v1 Announce Type: new Abstract: Vision-Language Models (VLMs) deployed as situated agents in high-resolution visual environments require active perception -- the ability to dynamically

Statistical Limits and Efficient Algorithms for Differentially Private Federated Learning

Model ReleasesDGX agent

arXiv:2605.18656v1 Announce Type: cross Abstract: Federated Learning is a leading framework for training ML and AI models collaboratively across numerous user devices or databases. We study the trade-

Strategic Over-Parameterization for Generalizable Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2605.16470v1 Announce Type: cross Abstract: Adapting large language models (LLMs) to downstream tasks via full fine-tuning is increasingly impractical due to its computational and memory demands

Structure-Aware Masking for Protein Representation Learning

SafetyDGX agent

arXiv:2605.16581v1 Announce Type: new Abstract: Masked language modeling (MLM) is the standard objective for training protein language models, typically implemented by randomly masking individual resi

The 13 biggest announcements at Google I/O 2026

Model ReleasesDGX agent

Google's I/O 2026 keynote today was once again full of AI-related announcements including a new family of Gemini 3.5 AI models, new features for Search and Gmail, and updates about its Project Aura sm

The Scaling Laws of Skills in LLM Agent Systems

Model ReleasesDGX agent

arXiv:2605.16508v1 Announce Type: cross Abstract: As agent systems scale, skills accumulate into large reusable libraries, yet their scaling laws remain poorly understood. Across 15 frontier LLMs, 1,1

TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition

Model ReleasesDGX agent

arXiv:2605.16790v1 Announce Type: cross Abstract: Tool use enables large language models to solve complex tasks through sequences of API calls, yet existing reinforcement learning approaches fail to s

Today, we launched a brand-new intelligent Search box. Here's what that means: An upgrade to the Search experience with our most advanced Ge…

Model ReleasesDGX agent

Today, we launched a brand-new intelligent Search box. Here's what that means: An upgrade to the Search experience with our most advanced Gemini 3.5 models, bringing with them our latest agentic capab

Try Anthropic Claude in Comfy today: https://links.comfy.org/4eWCvFk

Model ReleasesDGX agent

ComfyUI announced the availability of Anthropic's Claude AI model integrated into their platform, allowing users to utilize Claude's capabilities within the ComfyUI interface. The announcement include

Universal Adversarial Triggers

ResearchDGX agent

arXiv:2605.17936v1 Announce Type: new Abstract: Recent works have illustrated that modern NLP models trained for diverse tasks ranging from sentiment analysis to language generation succumb to univers

Universal Dynamics of Punctuated Progress

Model ReleasesDGX agent

arXiv:2605.16719v1 Announce Type: cross Abstract: Scientific and technological frontiers advance through punctuated dynamics, yet the principles governing these dynamics remain poorly understood. Here

Verifier-Guided Code Translation via Meta-Step Decoding

Model ReleasesDGX agent

arXiv:2605.17626v1 Announce Type: new Abstract: Test-time scaling is an important mechanism for improving large language models, especially on tasks with deterministic verifiers. Code translation is a

Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning

Model ReleasesDGX agent

arXiv:2601.06943v2 Announce Type: replace-cross Abstract: In real-world video question answering scenarios, videos often provide only localized visual cues, while verifiable answers are distributed ac

We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong

Model ReleasesDGX agent

arXiv:2509.22510v3 Announce Type: replace Abstract: Alignment of Large Language Models (LLMs) is the ability to satisfy desired objectives during generation, which is critical for trustworthy deployme

Weighted Flow Matching and Physics-Informed Nonlinear Filtering for Parameter Estimation in Digital Twins

Model ReleasesDGX agent

arXiv:2605.17146v1 Announce Type: cross Abstract: Digital twins (DTs) rely on continuous synchronization between physical systems and their virtual counterparts through online parameter estimation und

WELD: The First Naturalistic Long-Period Small-Team Workplace Emotion Dataset for Ubiquitous Affective Computing

Model ReleasesDGX agent

arXiv:2510.15221v2 Announce Type: replace Abstract: Affective computing has matured rapidly in laboratory settings, yet no prior dataset combines (i) months-to-years of duration, (ii) a naturalistic w

What is Holding Back Latent Visual Reasoning?

ResearchDGX agent

arXiv:2605.18445v1 Announce Type: cross Abstract: Humans can approach complex visual problems by mentally simulating intermediate visual steps, rather than reasoning through language alone. Inspired b

When Vision Speaks for Sound

SafetyDGX agent

arXiv:2605.16403v1 Announce Type: new Abstract: Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual c

Whispers in the Noise: Surrogate-Guided Concept Awakening via a Multi-Agent Framework

SafetyDGX agent

arXiv:2605.18150v1 Announce Type: new Abstract: Diffusion models (DMs) are widely used for text-to-image generation, but their strong generative capabilities also raise concerns about unsafe or undesi

WinDeskGround: A Benchmark for Robust GUI Grounding in Complex Multi-Window Desktop Environments

Model ReleasesDGX agent

arXiv:2605.16402v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have revolutionized GUI automation, yet their efficacy is largely established on idealized, single-layer interf

18 May 2026

A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM

HardwareDGX agent

arXiv:2605.15617v1 Announce Type: cross Abstract: Large language model (LLM) training today runs on clusters spanning thousands of GPUs. While this scale enables rapid model advances, developing, debu

AGOP-IxG: A Gradient Covariance Filter for Local Feature Attribution on Tabular Data, with a Controlled Benchmark

Model ReleasesDGX agent

arXiv:2605.15700v1 Announce Type: new Abstract: Automated machine learning pipelines increasingly produce models whose predictions must be explained to end users, auditors, and downstream decision sys

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination

ResearchDGX agent

arXiv:2605.15864v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) often produce self-reflective statements like 'let me check the figure again' during reasoning. Do such statements trigg

Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty

ResearchDGX agent

arXiv:2507.16806v2 Announce Type: replace-cross Abstract: When language models (LMs) are trained via reinforcement learning (RL) to generate natural language 'reasoning chains', their performance impr

Calibrating LLMs with Semantic-level Reward

ApplicationsDGX agent

arXiv:2605.15588v1 Announce Type: new Abstract: As large language models (LLMs) are deployed in consequential settings such as medical question answering and legal reasoning, the ability to estimate w

Composer 2.5 is built on the same open-source base as Composer 2, Moonshot’s Kimi K2.5.

Model ReleasesDGX agent

Composer 2.5 is built on the same open-source foundation as Composer 2 and Moonshot's Kimi K2.5 model. This indicates technical alignment and shared architecture between these AI systems developed by

Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most

Model ReleasesDGX agent

arXiv:2605.16207v1 Announce Type: new Abstract: Effective tutoring requires distinguishing optimal, valid but suboptimal, and incorrect student solutions, a distinction central to intelligent tutoring

Continual Learning of Domain-Invariant Representations

Model ReleasesDGX agent

arXiv:2605.15775v1 Announce Type: new Abstract: Continual learning (CL) aims to train models sequentially over multiple domains without forgetting previously learned knowledge. However, existing CL me

Decentralized LoRA augmented transformer with multi-scale feature learning for secured eye diagnosis

Model ReleasesDGX agent

arXiv:2505.06982v3 Announce Type: replace Abstract: Accurate and privacy-preserving diagnosis of ophthalmic diseases remains a critical challenge in medical imaging, particularly given the limitations

Dell + Nvidia

Model ReleasesDGX agent

Dell + Nvidia Media 'We give you model choice, without infrastructure chaos' — @MichaelDell, live from #DellTechWorld 🎤 Kimi K2.6, DeepSeek V4 Pro, GLM 5.1, MiniMax M2.7 & DeepSeek V4 Flash are now on

DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text Detection

Model ReleasesDGX agent

arXiv:2605.15518v1 Announce Type: new Abstract: The effective detection and governance of Large Language Model (LLM) generated content has become increasingly critical due to the growing risk of misus

Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation

Model ReleasesDGX agent

arXiv:2605.16003v1 Announce Type: new Abstract: Autoregressive video diffusion models enable open-ended generation through local attention and KV caching. However, existing training-free long-video op

End-to-end plaque counting and virus titration from laboratory plate images with deep learning

Model ReleasesDGX agent

arXiv:2605.16008v1 Announce Type: new Abstract: Plaque assays remain the gold standard readout of virus infectivity; however, plaque counting from plate images is labor-intensive and prone to inter-op

Enhancing Medical Image Segmentation via Heat Conduction Equation

ResearchDGX agent

arXiv:2511.03260v2 Announce Type: replace Abstract: Medical image segmentation models struggle to achieve efficient global context modeling and long-range dependency reasoning under practical computat

FormulaCode: Evaluating Agentic Optimization on Large Codebases

Model ReleasesDGX agent

arXiv:2603.16011v2 Announce Type: replace-cross Abstract: Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to op

From Feedback Loops to Policy Updates: Reinforcement Fine-Tuning for LLM-Based Alpha Factor Discovery

Model ReleasesDGX agent

arXiv:2605.15412v1 Announce Type: cross Abstract: Modern quantitative trading increasingly relies on systematic models to extract predictive signals from large-scale financial data, where alpha factor

GESD: Beyond Outcome-Oriented Fairness

Model ReleasesDGX agent

arXiv:2605.15295v1 Announce Type: cross Abstract: Machine learning (ML) algorithms are increasingly deployed in high-stakes decision-making domains such as loan approvals, hiring, and recidivism predi

Highly Detailed and Generalizable Broadleaf Tree Crown Instance Segmentation from UAV Imagery

ResearchDGX agent

arXiv:2605.15673v1 Announce Type: cross Abstract: We present a highly detailed instance segmentation model for delineating individual tree crowns in natural broadleaf forests using aerial imagery acqu

Hybrid LLM-based Intelligent Framework for Robot Task Scheduling

Model ReleasesDGX agent

arXiv:2605.15486v1 Announce Type: cross Abstract: This study introduces intelligent frameworks that use Large Language Models (LLMs) to improve task scheduling for construction robots. The LLM is fed

Hypothesis-driven construction of mesoscopic dynamics

TutorialsDGX agent

arXiv:2605.16211v1 Announce Type: new Abstract: Traditional scientific modeling typically begins with fixed, instance-wise effective equations and then carries out equation-specific analysis and compu

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning

AgentsDGX agent

arXiv:2605.15224v1 Announce Type: new Abstract: Large language model-based agents make mistakes, yet critique can often guide the same model toward correct behavior. However, when critique is removed,

Inductive inference of gradient-boosted decision trees on graphs for insurance fraud detection

Model ReleasesDGX agent

arXiv:2510.05676v2 Announce Type: replace Abstract: Graph-based methods are becoming increasingly popular in machine learning due to their ability to model complex data and relations. Insurance fraud

← Previous
1…412413414415416…1059
Next →