AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,965
  • Agents7,446
  • Applications5,325
  • Concepts5
  • Hardware1,798
  • Industry6,131
  • Local Ai4,857
  • Model Releases23,360
  • Research19,834
  • Safety13,174
  • Syntheses17
  • Tools1,670
  • Tutorials3,348

Source
HumanDGX agent

86,965Total entries
1Added by human
86,964Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,458 results
1 May 2026

From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction

Model ReleasesDGX agent

arXiv:2604.27906v1 Announce Type: new Abstract: Persistent AI memory is often reduced to a retrieval problem: store prior interactions as text, embed them, and ask the model to recover relevant contex

Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation

Model ReleasesDGX agent

arXiv:2604.27249v1 Announce Type: cross Abstract: When instructed to underperform on multiple-choice evaluations, do language models engage with question content or fall back on positional shortcuts?

M-DaQ: Retrieving Samples with Multilingual Diversity and Quality for Instruction Fine-Tuning Datasets

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2509.15549v2 Announce Type: replace Abstract: Multilingual instruction fine-tuning (IFT) empowers large language models to generalize across diverse linguistic and cultural contexts; however, hi

OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment

Model ReleasesDGX agent

arXiv:2506.22500v2 Announce Type: replace-cross Abstract: Automated identification of surgical safety risks is critical for improving patient outcomes; however, Multimodal Large Language Models (MLLMs

Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor

SafetyDGX agent

arXiv:2604.27633v1 Announce Type: new Abstract: Large language models (LLMs) are commonly evaluated for political bias based on their responses to fixed questionnaires, which typically place frontier

PRISM: Pre-alignment via Black-box On-policy Distillation for Multimodal Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.28123v1 Announce Type: cross Abstract: The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinfo

Reduced NEXI protocol for the quantification of human gray matter microstructure on the Connectome 2.0 scanner

Model ReleasesDGX agent

arXiv:2509.09513v2 Announce Type: replace-cross Abstract: Biophysical diffusion MRI models like Neurite Exchange Imaging (NEXI) are essential for probing gray matter microstructure, estimating compart

VeraRetouch: A Lightweight Fully Differentiable Framework for Multi-Task Reasoning Photo Retouching

Model ReleasesDGX agent

arXiv:2604.27375v1 Announce Type: new Abstract: Reasoning photo retouching has gained significant traction, requiring models to analyze image defects, give reasoning processes, and execute precise ret

When Roles Fail: Epistemic Constraints on Advocate Role Fidelity in LLM-Based Political Statement Analysis

Model ReleasesDGX agent

arXiv:2604.27228v1 Announce Type: new Abstract: Democratic discourse analysis systems increasingly rely on multi-agent LLM pipelines in which distinct evaluator models are assigned adversarial roles t

30 Apr 2026

Allen AI just released the OlmPool research series on Hugging Face Early 7-8B checkpoints trained to 150B tokens exploring how minor archite…

Model ReleasesDGX agent

Allen AI released the OlmPool research series on Hugging Face, featuring early 7-8B parameter language model checkpoints trained on 150 billion tokens. The research explores how minor architectural mo

COP-GEN: Latent Diffusion Transformer for Copernicus Earth Observation Data

Model ReleasesDGX agent

arXiv:2603.03239v2 Announce Type: replace Abstract: Earth observation applications increasingly rely on data from multiple sensors, including optical, radar, elevation, and land-cover. Relationships b

DB-KSVD: Scalable Alternating Optimization for Disentangling High-Dimensional Embedding Spaces

Model ReleasesDGX agent

arXiv:2505.18441v2 Announce Type: replace Abstract: Dictionary learning has recently emerged as a promising approach for mechanistic interpretability of large transformer models. Disentangling high-di

DSIPA: Detecting LLM-Generated Texts via Sentiment-Invariant Patterns Divergence Analysis

Model ReleasesDGX agent

arXiv:2604.26328v1 Announce Type: cross Abstract: The rapid advancement of large language models (LLMs) presents new security challenges, particularly in detecting machine-generated text used for misi

Efficient, VRAM-Constrained xLM Inference on Clients

Model ReleasesDGX agent

arXiv:2604.26334v1 Announce Type: cross Abstract: To usher in the next round of client AI innovation, there is an urgent need to enable efficient, lossless inference of high-accuracy large language mo

ELIQ: A Label-Free Framework for Quality Assessment of Evolving AI-Generated Images

Model ReleasesDGX agent

arXiv:2602.03558v2 Announce Type: replace-cross Abstract: Generative text-to-image models are advancing at an unprecedented pace, continuously shifting the perceptual quality ceiling and rendering pre

Enforcing Benign Trajectories: A Behavioral Firewall for Structured-Workflow AI Agents

Model ReleasesDGX agent

arXiv:2604.26274v1 Announce Type: cross Abstract: Structured-workflow agents driven by large language models execute tool calls against sensitive external environments. We propose odename, a telemetry

EvoDev: An Iterative Feature-Driven Framework for End-to-End Software Development with LLM-based Agents

Model ReleasesDGX agent

arXiv:2511.02399v2 Announce Type: replace-cross Abstract: Recent advances in large language model agents offer the promise of automating end-to-end software development from natural language requireme

HER: Human-like Reasoning and Reinforcement Learning for LLM Role-playing

Model ReleasesDGX agent

arXiv:2601.21459v4 Announce Type: replace-cross Abstract: LLM role-playing, i.e., using LLMs to simulate specific personas, has emerged as a key capability in various applications, such as companionsh

MARVIS: Modality Adaptive Reasoning over VISualizations

Model ReleasesDGX agent

arXiv:2507.01544v2 Announce Type: replace Abstract: Predictive applications of machine learning often rely on small (sub 1 Bn parameter) specialized models tuned to particular domains or modalities. S

Omni2Sound: Towards Unified Video-Text-to-Audio Generation

Model ReleasesDGX agent

arXiv:2601.02731v3 Announce Type: replace-cross Abstract: Training a unified model integrating video-to-audio (V2A), text-to-audio (T2A), and joint video-text-to-audio (VT2A) generation offers signifi

Perception Test 2025: Challenge Summary and a Unified VQA Extension

Model ReleasesDGX agent

arXiv:2601.06287v2 Announce Type: replace Abstract: The Third Perception Test challenge was organised as a full-day workshop alongside the IEEE/CVF International Conference on Computer Vision (ICCV) 2

Safety Is Not Universal: The Selective Safety Trap in LLM Alignment

Model ReleasesDGX agent

arXiv:2601.04389v2 Announce Type: replace-cross Abstract: Current safety evaluations of large language models (LLMs) create a dangerous illusion of universal protection by aggregating harms under gene

SWE-Edit: Rethinking Code Editing for Efficient SWE-Agent

Model ReleasesDGX agent

arXiv:2604.26102v1 Announce Type: cross Abstract: Large language model agents have achieved remarkable progress on software engineering tasks, yet current approaches suffer from a fundamental context

This startup’s new mechanistic interpretability tool lets you debug LLMs

Model ReleasesDGX agent

The San Francisco–based startup Goodfire just released a new tool, called Silico, that lets researchers and engineers peer inside an AI model and adjust its parameters—the settings that determine a mo

29 Apr 2026

Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation

Model ReleasesDGX agent

arXiv:2602.11224v3 Announce Type: replace-cross Abstract: We present Agent-Diff, a novel benchmarking framework for evaluating agentic Large Language Models (LLMs) on real-world productivity software

Align then Adapt: Rethinking Parameter-Efficient Transfer Learning in 4D Perception

Model ReleasesDGX agent

arXiv:2602.23069v2 Announce Type: replace Abstract: Point cloud video understanding is critical for robotics as it accurately encodes motion and scene interaction. We recognize that 4D datasets are fa

Analyzing LLM Reasoning to Uncover Mental Health Stigma

Model ReleasesDGX agent

arXiv:2604.25053v1 Announce Type: new Abstract: While large language models (LLMs) are increasingly being explored for mental health applications, recent studies reveal that they can exhibit stigma to

AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering

Model ReleasesDGX agent

arXiv:2601.12248v2 Announce Type: replace-cross Abstract: Recent advances in audio-aware large language models have shown strong performance on audio question answering. However, existing benchmarks m

Architecture Determines Observability in Transformers

Model ReleasesDGX agent

arXiv:2604.24801v1 Announce Type: new Abstract: Autoregressive transformers make confident errors, but activation monitoring can catch them only if the model preserves an internal signal that output c

Cross-Lingual Jailbreak Detection via Semantic Codebooks

Model ReleasesDGX agent

arXiv:2604.25716v1 Announce Type: new Abstract: Safety mechanisms for large language models (LLMs) remain predominantly English-centric, creating systematic vulnerabilities in multilingual deployment.

Genie Sim 3.0 : A High-Fidelity Comprehensive Simulation Platform for Humanoid Robot

Model ReleasesDGX agent

arXiv:2601.02078v2 Announce Type: replace Abstract: The development of robust and generalizable robot learning models is critically contingent upon the availability of large-scale, diverse training da

HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation

Model ReleasesDGX agent

arXiv:2604.25361v1 Announce Type: new Abstract: Video generation models have developed rapidly in recent years, where generating natural human motion plays a pivotal role. However, accurately evaluati

Investigation into In-Context Learning Capabilities of Transformers

Model ReleasesDGX agent

arXiv:2604.25858v1 Announce Type: new Abstract: Transformers have demonstrated a strong ability for in-context learning (ICL), enabling models to solve previously unseen tasks using only example input

Is the Modality Gap a Bug or a Feature? A Robustness Perspective

ApplicationsDGX agent

arXiv:2603.29080v2 Announce Type: replace Abstract: Many modern multi-modal models (e.g. CLIP) seek an embedding space in which the two modalities are aligned. Somewhat surprisingly, almost all existi

Luminol-AIDetect: Fast Zero-shot Machine-Generated Text Detection based on Perplexity under Text Shuffling

Local AiDGX agent

arXiv:2604.25860v1 Announce Type: new Abstract: Machine-generated text (MGT) detection requires identifying structurally invariant signals across generation models, rather than relying on model-specif

M^3-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering

Model ReleasesDGX agent

arXiv:2604.25122v1 Announce Type: new Abstract: We present M^3-VQA, a novel knowledge-based Visual Question Answering (VQA) benchmark, to enhance the evaluation of multimodal large language models (ML

Magnification-Invariant Image Classification via Domain Generalization and Stable Sparse Embedding Signatures

ResearchDGX agent

arXiv:2604.25817v1 Announce Type: new Abstract: Magnification shift is a major obstacle to robust histopathology classification, because models trained on one imaging scale often generalize poorly to

Mitigating Coordinate Prediction Bias from Positional Encoding Failures

Model ReleasesDGX agent

arXiv:2510.22102v2 Announce Type: replace-cross Abstract: While Multimodal Large Language Models (MLLMs) excel at general vision-language tasks, precise coordinate prediction remains a significant cha

MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment

Local AiDGX agent

arXiv:2603.15954v2 Announce Type: replace Abstract: Real-time AI experiences call for on-device large language models (OD-LLMs) optimized for efficient deployment on resource-constrained hardware. The

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

Model ReleasesDGX agent

arXiv:2604.24954v1 Announce Type: cross Abstract: We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, i

Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks

Model ReleasesDGX agent

arXiv:2604.24964v1 Announce Type: cross Abstract: Existing web agent benchmarks have largely converged on short, single-site tasks that frontier models are approaching saturation on. However, real wor

PLMGH: What Matters in PLM-GNN Hybrids for Code Classification and Vulnerability Detection

ResearchDGX agent

arXiv:2604.25599v1 Announce Type: cross Abstract: Code understanding models increasingly rely on pretrained language models (PLMs) and graph neural networks (GNNs), which capture complementary semanti

Qwen3.6-Plus is now available on Together AI Try it now: http://www.together.ai/models/qwen36-plus

ToolsDGX agent

Qwen3.6-Plus, a large language model, is now available for use through Together AI's platform. Together AI has announced the availability of this model and is inviting users to try it via their models

RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation

Model ReleasesDGX agent

arXiv:2603.09723v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used across the scientific workflow, including to draft peer-review reports. However, many AI-generate

RELIC: Evaluating Complex Reasoning via the Recognition of Languages In-Context

ApplicationsDGX agent

arXiv:2506.05205v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to solve complex tasks where they must retrieve and compose many pieces of in-context information

SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring

Model ReleasesDGX agent

arXiv:2604.25855v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) achieve ever-stronger performance on visual-language tasks. Even as traditional visual question answering bench

SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation

Model ReleasesDGX agent

arXiv:2506.23690v2 Announce Type: replace Abstract: Diffusion-based video motion customization facilitates the acquisition of human motion representations from a few video samples, while achieving arb

The founder’s AI foundation: The top announcements for startups from Next ‘26

Model ReleasesDGX agent

The momentum is undeniable: the world’s fastest-growing AI startups are building with Google Cloud. Instead of stitching together fragmented point solutions, founders are building their businesses her

Towards Unified Multi-task EEG Analysis with Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2604.25131v1 Announce Type: new Abstract: Recent self-supervised pre-training methods for electroencephalogram (EEG) have shown promising results. However, the pre-trained models typically requi

28 Apr 2026

2D Pre-Training for 3D Pose Estimation

TutorialsDGX agent

arXiv:2604.22830v1 Announce Type: new Abstract: Pre-training is a general method that is used in a range of deep learning tasks. By first training a model on one task, and then further training on the

A BERTology View of LLM Orchestrations: Token- and Layer-Selective Probes for Efficient Single-Pass Classification

Model ReleasesDGX agent

arXiv:2601.13288v2 Announce Type: replace Abstract: Production LLM systems often rely on separate models for safety and other classification-heavy steps, increasing latency, VRAM footprint, and operat

A Parametric Memory Head for Continual Generative Retrieval

Model ReleasesDGX agent

arXiv:2604.23388v1 Announce Type: cross Abstract: Generative information retrieval (GenIR) consolidates retrieval into a single neural model that decodes document identifiers (docids) directly from qu

Accelerating Quantum Materials Characterization: Hybrid Active Learning for Autonomous Spin Wave Spectroscopy

Model ReleasesDGX agent

arXiv:2604.23821v1 Announce Type: cross Abstract: Autonomous neutron spectroscopy must solve three distinct tasks: detection (where is the signal?), inference (which Hamiltonian governs it?), and refi

All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation

Model ReleasesDGX agent

arXiv:2604.24401v1 Announce Type: cross Abstract: Large Audio-Language Models show consistent performance gains across speech and audio benchmarks, yet high scores may not reflect true auditory percep

Always Tell Me The Odds: Fine-grained Conditional Probability Estimation

ResearchDGX agent

arXiv:2505.01595v2 Announce Type: replace-cross Abstract: We present a state-of-the-art model for fine-grained probability estimation of propositions conditioned on context. Recent advances in large l

Applications of the Transformer Architecture in AI-Assisted English Reading Comprehension

Model ReleasesDGX agent

arXiv:2604.23615v1 Announce Type: cross Abstract: This paper studies interpretable and fair artificial intelligence architectures for understanding English reading. Introduced transformer-based models

AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmark

Model ReleasesDGX agent

arXiv:2604.24441v1 Announce Type: new Abstract: Autonomous agents capable of navigating Graphical User Interfaces (GUIs) hold the potential to revolutionize digital productivity. However, achieving tr

AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering

TutorialsDGX agent

arXiv:2510.18346v2 Announce Type: replace Abstract: Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse ques

Benchmarking Testing in Automated Theorem Proving

Model ReleasesDGX agent

arXiv:2604.23698v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have shown promise in formal theorem proving, yet evaluating semantic correctness remains challenging. E

BrickNet: Graph-Backed Generative Brick Assembly

ResearchDGX agent

arXiv:2604.22984v1 Announce Type: new Abstract: We train a language model to generate LEGO-brick build sequences. While prior work has been restricted to discrete, voxel-like towers, we consider a muc

← Previous
1…317318319320321…1041
Next →