AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,611 results
5 Jun 2026

Global Cross-Modal Geo-Localization: A Million-Scale Dataset and a Physical Consistency Learning Framework

Model ReleasesDGX agent

arXiv:2603.08491v2 Announce Type: replace Abstract: Cross-modal Geo-localization (CMGL) matches ground-level text descriptions with geo-tagged aerial imagery, which is crucial for pedestrian navigatio

Google Cloud says its SpaceX compute deal is a 'short-term' agreement 'to ensure we have bridge capacity to meet surging customer demand' for Gemini Enterprise (Kate Conger/New York Times)

Model ReleasesDGX agent

Kate Conger / New York Times: Google Cloud says its SpaceX compute deal is a “short-term” agreement “to ensure we have bridge capacity to meet surging customer demand” for Gemini Enterprise — Elon Mus

Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2511.20158v2 Announce Type: replace Abstract: While continual visual instruction tuning (CVIT) has shown promise in adapting multimodal large language models (MLLMs), existing studies predominan

Harnessing Structural Context for Entity Alignment Foundation Models

Model ReleasesDGX agent

arXiv:2606.06109v1 Announce Type: new Abstract: Entity alignment (EA) aims to identify equivalent entities across heterogeneous knowledge graphs (KGs) and is a key component of knowledge fusion and cr

Here’s this week’s shipping recap 👇 — Nano Banana 2 & Nano Banana Pro are now GA and available via the Gemini Enterprise Agent Platform, Ge…

Model ReleasesDGX agent

Here’s this week’s shipping recap 👇 — Nano Banana 2 & Nano Banana Pro are now GA and available via the Gemini Enterprise Agent Platform, Gemini API, and in @GoogleAIStudio —Co-Scientist, our new multi

HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps

Model ReleasesDGX agent

arXiv:2601.02730v3 Announce Type: replace Abstract: Visual localization on standard-definition (SD) maps has emerged as a promising low-cost and scalable solution for autonomous driving. However, exis

Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

Model ReleasesDGX agent

arXiv:2606.06388v1 Announce Type: cross Abstract: Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly pos

I always appreciate the opportunity to discuss @LawZero_ and our approach to honest, reliable AI. Working on the Scientist AI with my brilli…

Model ReleasesDGX agent

I always appreciate the opportunity to discuss @LawZero_ and our approach to honest, reliable AI. Working on the Scientist AI with my brilliant colleagues at LawZero has made me very confident that we

IA-RAG: Interval-Algebra-Driven Temporal Reasoning for Dynamic Knowledge Retrieval

Model ReleasesDGX agent

arXiv:2606.06044v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has shown strong effectiveness in grounding Large Language Models (LLMs) with external knowledge. However, existing

If Claude is good enough for Nobel Prize winners it is good enough for you https://arxiv.org/abs/2606.03300

Model ReleasesDGX agent

This post appears to reference Claude AI's capabilities and performance, likely highlighting how the model has been used or endorsed by notable researchers or Nobel Prize winners to establish credibil

Improving Answer Extraction in Context-based Question Answering Systems Using LLMs

Model ReleasesDGX agent

arXiv:2606.06197v1 Announce Type: new Abstract: Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs). However, they still face challenges in a

Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing

Model ReleasesDGX agent

arXiv:2606.05172v1 Announce Type: cross Abstract: Diffusion-based image editing has achieved strong visual fidelity under natural language instructions, yet most existing systems still operate at the

KV-Control: Parameter-Efficient K/V Injection for Trajectory-Controlled Text-to-Motion

Model ReleasesDGX agent

arXiv:2606.05624v1 Announce Type: new Abstract: Text-conditioned 3D human motion models now synthesize plausible motions from prompts, but practical animation and embodied-agent workflows rarely stop

LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents

Model ReleasesDGX agent

arXiv:2606.06087v1 Announce Type: new Abstract: Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substa

Less is MoE: Trimming Experts in Domain-Specialist Language Models

Model ReleasesDGX agent

arXiv:2606.05538v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models achieve strong performance through conditional computation, but their large parameter footprint poses deployment chall

LiAuto-GeoX: Efficient Grounded Driving Transformer

Model ReleasesDGX agent

arXiv:2606.05774v1 Announce Type: new Abstract: Dense 3D reconstruction has demonstrated immense potential for spatial understanding, yet its viability as a real-time, onboard representation for auton

LightVesselNet: An Ultra-Lightweight Sub-100K Parameter Network for Retinal Blood Vessel Segmentation

Model ReleasesDGX agent

arXiv:2606.05354v1 Announce Type: new Abstract: Retinal blood vessel segmentation plays a vital role in the early detection of diabetic retinopathy and glaucoma. While recent deep learning models have

LLM-Guided ANN Index Optimization for Human-Object Interaction Retrieval

Model ReleasesDGX agent

arXiv:2606.05489v1 Announce Type: new Abstract: Retrieval systems underpin modern AI applications -- spanning visual search, recommendation engines, and multi-modal question answering. Modern multi-st

Localizing Prompt Ambiguity in Large Language Models with Probe-Targeted Attribution

Model ReleasesDGX agent

arXiv:2606.05486v1 Announce Type: new Abstract: Prompt ambiguity is a common source of failure in large language models, but is difficult to localize because it is a latent property of the prompt, whi

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

Model ReleasesDGX agent

arXiv:2606.05677v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have advanced image and video understanding and can increasingly handle longer visual inputs. Long-horizon ta

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

Model ReleasesDGX agent

arXiv:2606.06042v1 Announce Type: new Abstract: Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier fie

LoRi: Low-Rank Distillation for Implicit Reasoning

Model ReleasesDGX agent

arXiv:2606.05315v1 Announce Type: new Abstract: Implicit chain-of-thought (iCoT) methods aim to internalize reasoning in large language models, but often underperform explicit CoT prompting. We empiri

MAviS: A Multimodal Conversational Assistant For Avian Species

Model ReleasesDGX agent

arXiv:2603.07294v2 Announce Type: replace Abstract: Fine-grained understanding and species-specific multimodal question answering are vital for advancing biodiversity conservation and ecological monit

MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

Model ReleasesDGX agent

arXiv:2606.05177v1 Announce Type: new Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and

MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following

Model ReleasesDGX agent

arXiv:2606.06058v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards is ideal for multi-constraint instruction following, yet standard group-relative policy optimization (G

Nemotron 3 Ultra is now available for Pro and Max subscribers on Perplexity and Computer. It's @nvidia's new open model built for long-runni…

Model ReleasesDGX agent

Nvidia's Nemotron 3 Ultra, an open-source model designed for long-context tasks, is now available to Pro and Max subscribers on Perplexity and Computer. This new model represents Nvidia's latest contr

Noise-Aware Visual Representation Learning for Medical Visual Question Answering

Model ReleasesDGX agent

arXiv:2606.05535v1 Announce Type: new Abstract: Medical visual question answering (Med-VQA) has strong potential for clinical decision support by enabling AI models to interpret medical images and ans

Oklch+: A Three-Parameter Extension of Oklab for Improved Color Difference Prediction

Model ReleasesDGX agent

arXiv:2606.05255v1 Announce Type: cross Abstract: Oklab and its cylindrical representation Oklch are widely adopted in interpolation and design workflows as perceptually motivated color spaces, but th

OLIVE: Online Low-Rank Incremental Learning for Efficient Adaptive Exoskeletons

Model ReleasesDGX agent

arXiv:2606.05234v1 Announce Type: new Abstract: Wearable exoskeleton systems hold promise for restoring mobility in individuals with physical impairments, yet most existing controllers rely on static

Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection

Model ReleasesDGX agent

arXiv:2606.06481v1 Announce Type: new Abstract: As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-writt

PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis

Model ReleasesDGX agent

arXiv:2606.05176v1 Announce Type: new Abstract: While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptation to domain-s

Physics-Guided Deep Unfolding for Blind Cross-Sensor Spectral Super-Resolution via Learning the Spectral Transformation Function

Model ReleasesDGX agent

arXiv:2606.05759v1 Announce Type: new Abstract: Hyperspectral imaging provides rich spectral information for quantitative remote sensing, yet hyperspectral sensors remain costly and thus unavailable i

PlanBench-V: A Spatial Planning Map Benchmark for Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.05744v1 Announce Type: new Abstract: Spatial planning maps are central to territorial governance, translating planning objectives, regulations, and spatial strategies into visual forms for

Predict and Reconstruct: Joint Objectives for Self-Supervised Language Representation Learning

Model ReleasesDGX agent

arXiv:2606.05173v1 Announce Type: new Abstract: Masked language modelling (MLM) has been the dominant pre-training objective for text encoders since BERT, yet it encourages representations that are st

ProSPy: A Profiling-Driven SQL-Python Agentic Framework for Enterprise Text-to-SQL

Model ReleasesDGX agent

arXiv:2606.05836v1 Announce Type: new Abstract: Large language models have substantially advanced Text-to-SQL systems, yet applying them to enterprise-scale databases remains challenging. Real-world d

RAPTOR+: A Visually Grounded Vision-Language Framework to Improve Clinical Trust and Auditability in Automated Cancer Referral Processing

Model ReleasesDGX agent

arXiv:2605.25956v2 Announce Type: replace Abstract: Urgent suspected colorectal cancer (CRC) referrals create operational bottlenecks because semi-structured clinical documents often require manual re

ReasoningFlow: Discourse Structures for Understanding LLM Reasoning Traces

Model ReleasesDGX agent

arXiv:2606.05402v1 Announce Type: new Abstract: Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluat

RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit

Model ReleasesDGX agent

arXiv:2606.06027v1 Announce Type: cross Abstract: Community-conditioned language model adaptation requires choices about data collection, community definition, and evaluation that are currently made i

Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)

Model ReleasesDGX agent

arXiv:2606.05901v1 Announce Type: new Abstract: Large language models (LLMs) have fundamentally transformed the landscape of Natural Language Processing. Despite these advances, LLMs and LLM-based sys

Representing Research Attention as Contextually Structured Flows

Model ReleasesDGX agent

arXiv:2606.05895v1 Announce Type: new Abstract: Research attention is widely used as an indicator of visibility, influence, and societal uptake, yet it is typically represented as aggregated counts th

Rethinking LoRA Memory Through the Lens of KV Cache Compression

Model ReleasesDGX agent

arXiv:2606.05698v1 Announce Type: new Abstract: Parametric retrieval augmentation encodes document information into lightweight, document-specific modules such as LoRA adapters, reducing the need to i

Safe Embodied AI for Long-horizon Tasks: A Cross-layer Analysis of Robotic Manipulation

Model ReleasesDGX agent

arXiv:2606.05660v1 Announce Type: new Abstract: Embodied AI systems are increasingly expected to reason and act over extended horizons in physical environments. This growing capability brings safety t

Scaffold, Not Vocabulary? A Controlled, Two-Tier, Pre-Registered Study of a Popperian Code-Generation Skill

Model ReleasesDGX agent

arXiv:2606.06454v1 Announce Type: cross Abstract: Large language models increasingly write, review, and judge code, and a fast-growing practice equips them with prompt 'skills' that ask the model to r

Seeing is Believing? Evaluating Vision-Language Model Susceptibility in Agent-to-Agent Multimodal Persuasion

Model ReleasesDGX agent

arXiv:2510.22768v2 Announce Type: replace Abstract: As autonomous agents increasingly interact, they inevitably attempt to influence one another. While prior work in text-only settings has explored th

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.05702v1 Announce Type: cross Abstract: Recent advancements in Vision-Language Models (VLMs) have significantly enhanced their ability to interpret complex visual semantics, yet their capaci

Self-supervised User Profile Generation for Personalization

Model ReleasesDGX agent

arXiv:2606.05336v1 Announce Type: new Abstract: Personalizing large language models (LLMs) has become a central challenge as LLMs are deployed across recommendation, search, dialogue, and content gene

ShotCrop^3: Cropping Human-Centric Images into Cinematic Triple-Shot Compositions

Model ReleasesDGX agent

arXiv:2606.05635v1 Announce Type: new Abstract: Prior work on aesthetic composition typically produces a single aesthetically pleasing crop, overlooking the narrative value of composing multiple shots

SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations

Model ReleasesDGX agent

arXiv:2606.05563v1 Announce Type: cross Abstract: Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and

Sources say xAI used Claude models for distillation and training, including using personal accounts and the intermediary service Blackbox AI after being cut off (Grace Kay/The Information)

Model ReleasesDGX agent

Grace Kay / The Information: Sources say xAI used Claude models for distillation and training, including using personal accounts and the intermediary service Blackbox AI after being cut off — SpaceX's

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges

Model ReleasesDGX agent

arXiv:2606.05384v1 Announce Type: cross Abstract: LLM-as-judge evaluation is widely used in benchmarking pipelines, where model outputs are compared and ranked using automated evaluators. These pipeli

Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference

Model ReleasesDGX agent

arXiv:2606.05308v1 Announce Type: cross Abstract: With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-la

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset

Model ReleasesDGX agent

arXiv:2606.06338v1 Announce Type: new Abstract: Video question answering (VideoQA) aims to answer questions about given videos. While existing approaches excel on factoid VideoQA, they struggle with d

SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents

Model ReleasesDGX agent

arXiv:2606.05761v1 Announce Type: cross Abstract: Persistent AI assistants, such as OpenClaw, accumulate large collections of related memories over long-term interactions. As these memories grow, they

TARPO: Token-Wise Latent-Explicit Reasoning via Action-Routing Policy Optimization

Model ReleasesDGX agent

arXiv:2606.05859v1 Announce Type: new Abstract: Latent reasoning has emerged as a promising alternative to discrete Chain-of-Thought (CoT) in large language models (LLMs), enabling more expressive rea

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison

Model ReleasesDGX agent

arXiv:2606.05436v1 Announce Type: cross Abstract: Summarizing the latest medical literature to guide clinical decision-making is essential for evidence-based medicine and high-quality patient care. Ye

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework

Model ReleasesDGX agent

arXiv:2606.05570v1 Announce Type: new Abstract: Repository-level coding benchmarks face a trade-off between task difficulty and evaluation reliability: tasks that challenge frontier models often invol

TextWand: A Unified Framework for Scene Text Editing

Model ReleasesDGX agent

arXiv:2606.05730v1 Announce Type: new Abstract: We propose TextWand, a general-purpose framework that unifies scene text removal, generation, and replacement into a single model. By decomposing comple

The Granularity Gap: A Multi-Dimensional Longitudinal Audit of Sycophancy in Gemini Models

Model ReleasesDGX agent

arXiv:2606.05183v1 Announce Type: new Abstract: Large language models are increasingly deployed as high-stakes advisors, yet standard alignment benchmarks treat sycophancy as a binary failure mode. We

The latest AI news we announced in May 2026

Model ReleasesDGX agent

Google's May 2026 AI updates center on the new 'agentic' era, featuring the Gemini 3.5 model and Gemini Omni for advanced reasoning and creation. Gemini Omni is a new model that can create anything fr

The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?

Model ReleasesDGX agent

arXiv:2504.10020v4 Announce Type: replace Abstract: Contrastive decoding strategies are widely used to reduce object hallucinations in multimodal large language models (MLLMs). These methods work by c

← Previous
1…167168169170171…377
Next →