AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,628 results
2 Jun 2026

TECCI: Tricky Edits of Collected and Curated Images

Model ReleasesDGX agent

arXiv:2606.01213v1 Announce Type: cross Abstract: Despite tremendous recent progress, current text-guided image editing methods still struggle with many aspects of editing involving instruction follow

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation

Model ReleasesDGX agent

arXiv:2606.01031v1 Announce Type: cross Abstract: Audio-driven talking-head generation has advanced rapidly, yet existing evaluation protocols mainly rely on frame-wise metrics that assume strict temp

TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.24470v2 Announce Type: replace Abstract: Video-text retrieval has witnessed remarkable progress driven by large-scale vision-language pretraining, yet most existing approaches inherit an im

TextFake: Benchmarking AI-Generated Image Detection on Text-Rich Images

Model ReleasesDGX agent

arXiv:2606.01050v1 Announce Type: new Abstract: Recent AI-generated image (AIGI) detectors perform well on natural-image benchmarks, but their behavior on text-rich forgeries, such as fabricated scree

The Assistant as a Privileged Persona: A canonical reference in cross-persona self-recognition

Model ReleasesDGX agent

arXiv:2606.00545v1 Announce Type: new Abstract: Post-trained language models can recognize their own outputs from a sentence or two out of context. In a companion paper itep{jack2026twomodes} we showe

The Case for Model Science: Verify, Explore, Steer, Refine

Model ReleasesDGX agent

arXiv:2606.01189v1 Announce Type: new Abstract: We argue that the AI community is now ready to move beyond benchmarking and consolidate scattered efforts in model analysis into a systematic discipline

The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing

Model ReleasesDGX agent

arXiv:2606.02184v1 Announce Type: cross Abstract: These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academ

The Image Reconstruction Game: Drawing Common Ground Through Iterative Multimodal Dialogue

Model ReleasesDGX agent

arXiv:2606.01901v1 Announce Type: cross Abstract: We introduce the Image Reconstruction Game, a fully automated benchmark in which a vision-language model issues corrective instructions to an image ge

The Invisible Coalition Partner: How LLMs Vote When Democracy Gets Concrete

Model ReleasesDGX agent

arXiv:2606.00048v1 Announce Type: cross Abstract: Prior research has established that instruction-tuned large language models exhibit left-of-center political bias, measured exclusively through abstra

The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models

Model ReleasesDGX agent

arXiv:2605.05427v2 Announce Type: replace Abstract: Refusal rates are a poor proxy for LLM safety, i.e., a model may over-refuse benign prompts while still complying with harmful ones. We audit both f

The Shape of Wisdom: Decision Trajectories in Language Models

Model ReleasesDGX agent

arXiv:2606.01202v1 Announce Type: new Abstract: Language models do not simply choose an answer at the output layer. In a 9,000-trajectory MMLU study across Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct,

This is actually one of the main advantages startups have over frontier labs, as long as there's a healthy spectrum of open-weight to closed…

Model ReleasesDGX agent

This is actually one of the main advantages startups have over frontier labs, as long as there's a healthy spectrum of open-weight to closed-weight models on the cost-performance curve. Building a mod

This is also available on the Claude Blog! https://claude.com/blog/a-harness-for-every-task-dynamic-workflows-in-claude-code

Model ReleasesDGX agent

This post discusses dynamic workflows in Claude Code, enabling flexible task automation and execution patterns. The content is available on the official Claude Blog and addresses how Claude can be har

🌞This is big Local AI news! A new open-source Computer-Use LLM has just launched. Holo 3.1 is H Company’s (🇫🇷) new local computer-use age…

Model ReleasesDGX agent

🌞This is big Local AI news! A new open-source Computer-Use LLM has just launched. Holo 3.1 is H Company’s (🇫🇷) new local computer-use agent model that beats Qwen3.5-397B, Kimi-K2.5, and Sonnet 4.6! Si

this is fine 🐶☕️🔥

Model ReleasesDGX agent

This post likely references the popular 'This is Fine' meme featuring a dog in a burning room, often used to comment on problematic situations being accepted or ignored. Without access to the specific

Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs

Model ReleasesDGX agent

arXiv:2606.00592v1 Announce Type: new Abstract: Effective visual communication stems from the harmony of multiple design principles, such as readability, contrast, alignment, overlap, and coherence, w

Time-Optimal Collision Avoidance Via a Greedy Polynomial Backward Sweep

Model ReleasesDGX agent

arXiv:2606.01169v1 Announce Type: cross Abstract: Spacecraft collision avoidance for low-thrust satellites often requires determining not only how to maneuver, but also how late a maneuver can begin w

TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning

Model ReleasesDGX agent

arXiv:2606.01498v1 Announce Type: cross Abstract: Time series data inform critical decisions across many real-world domains. While large language model (LLM) agents can analyze data through natural la

Tiny Recursive Models for Solving the J2-Perturbed Lambert Problem

Model ReleasesDGX agent

arXiv:2606.00895v1 Announce Type: cross Abstract: This paper presents a fast, recursive neural solver for the J2-perturbed Lambert problem based on Tiny Recursive Models (TRM), termed the TRM-Perturbe

TLG: Temporal-Logic Grounding for Video Question Answering via Source-Annotation Reconstruction and Category-Targeted Reasoning

Model ReleasesDGX agent

arXiv:2606.01591v1 Announce Type: new Abstract: The TimeLogic Challenge evaluates formal temporal-logic reasoning over video - 16 operators (before, after, until, since, always, co-occur, ordering, ..

Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners

Model ReleasesDGX agent

arXiv:2606.01810v1 Announce Type: new Abstract: Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning. Thi

Topological Ignorability for Structural Causal Effects Beyond Means

Model ReleasesDGX agent

arXiv:2606.01184v1 Announce Type: cross Abstract: Many interventions alter the structure of an outcome distribution rather than its mean: they can split a population into disconnected regimes, create

Toward accurate RUL and SoH estimation using reinforced graph-based physics-informed neural networks enhanced with dynamic weights

Model ReleasesDGX agent

arXiv:2507.09766v2 Announce Type: replace-cross Abstract: Accurate estimation of Remaining Useful Life (RUL) and State of Health (SoH) is essential for reliable Prognostics and Health Management (PHM)

Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification

Model ReleasesDGX agent

arXiv:2509.16635v2 Announce Type: replace Abstract: In real applications, person re-identification (ReID) is expected to retrieve the target person at any time, including both daytime and nighttime, r

Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models

Model ReleasesDGX agent

arXiv:2606.00919v1 Announce Type: new Abstract: Large language models (LLMs) have seen widespread adoption across various domains, yet their reliability is frequently undermined by hallucinations - re

Towards Multidisciplinary Summarization of Hospital Stays: Efficient Sentence-Level Clinical Provenance Categorization

Model ReleasesDGX agent

arXiv:2606.02487v1 Announce Type: new Abstract: Effective 'all-team' summarization in high-complexity settings like the Neonatal Intensive Care Unit (NICU) requires aggregating insights from diverse d

Towards Simple and Provable Parameter-Free Adaptive Gradient Methods

Model ReleasesDGX agent

arXiv:2412.19444v2 Announce Type: replace Abstract: Optimization algorithms such as AdaGrad and Adam have significantly advanced the training of deep models by dynamically adjusting the learning rate

TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

Model ReleasesDGX agent

arXiv:2606.01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existi

Tree-Structured Parzen Estimator: Understanding Its Algorithm Components and Their Roles for Better Empirical Performance

Model ReleasesDGX agent

arXiv:2304.11127v5 Announce Type: replace-cross Abstract: Recent scientific advances require complex experiment design, necessitating the meticulous tuning of many experiment parameters. Tree-structur

Trump signs executive order to review AI models before they’re released

Model ReleasesDGX agent

President Donald Trump signed an executive order Tuesday creating a 'voluntary framework' for AI companies to share their frontier models with the federal government before they're released 'to promot

TrustLDM: Benchmarking Trustworthiness in Language Diffusion Models

Model ReleasesDGX agent

arXiv:2606.00023v1 Announce Type: cross Abstract: The rapid development of Language Diffusion Models (LDMs) challenges the dominant position of auto-regressive competitors in language processing. Howe

Truth, Trust, and Trouble: Medical AI on the Edge

Model ReleasesDGX agent

arXiv:2507.02983v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) hold significant promise for transforming digital health by enabling automated medical question answering. Howeve

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment

Model ReleasesDGX agent

arXiv:2606.01456v1 Announce Type: cross Abstract: Large language models are increasingly deployed as advisors whose objective is not aligned with the user's: recommenders optimize for engagement, sale

TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages

Model ReleasesDGX agent

arXiv:2606.01322v1 Announce Type: cross Abstract: Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly African ones, c

Turning Back Without Forgetting: Selective Backward Refinement for Parameter-Efficient Continual Learning

Model ReleasesDGX agent

arXiv:2606.01379v1 Announce Type: new Abstract: While prompt-based parameter-efficient continual learning mitigates catastrophic forgetting by isolating task-specific prompts, this isolation also limi

TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation

Model ReleasesDGX agent

arXiv:2606.02320v1 Announce Type: new Abstract: Deep Research Agents have shown strong capability in multi-step information retrieval, reasoning, and long-form report generation, but existing benchmar

Uber says it has limited all employees to $1,500 in monthly token spending per AI coding tool 'to responsibly encourage agentic AI adoption' (Natalie Lung/Bloomberg)

Model ReleasesDGX agent

Natalie Lung / Bloomberg: Uber says it has limited all employees to $1,500 in monthly token spending per AI coding tool “to responsibly encourage agentic AI adoption” — Uber Technologies Inc. has set

Uncovering Competency Gaps in Large Language Models and Their Benchmarks

Model ReleasesDGX agent

arXiv:2512.20638v2 Announce Type: replace-cross Abstract: The evaluation of large language models relies heavily on standardized benchmarks. These benchmarks provide useful aggregated metrics, but can

Understanding Identity Continuity in Thermal Video through Scene-Level Consistency

Model ReleasesDGX agent

arXiv:2606.01694v1 Announce Type: cross Abstract: Thermal pedestrian MOT remains challenging because weak appearance cues and frequent detection interruptions cause severe trajectory fragmentation. We

Understanding LLM Behavior in Multi-Target Cross-Lingual Summarization

Model ReleasesDGX agent

arXiv:2606.01252v1 Announce Type: cross Abstract: Multi-target cross-lingual text summarization (MTXLS), which summarizes a source document into multiple target languages, is increasingly important as

UniD^3: A Knowledge Graph-Enhanced RAG Framework for Drug-Disease Discovery and Reasoning

Model ReleasesDGX agent

arXiv:2606.01394v1 Announce Type: new Abstract: Systematic characterization of drug-disease relationships is essential for drug discovery and repurposing, yet is hindered by the heterogeneity and rapi

Universal Quantum Transformer

Model ReleasesDGX agent

arXiv:2606.00045v1 Announce Type: new Abstract: Classical continuous-space neural networks fundamentally struggle to lock into exact mathematical symmetries, such as modular arithmetic and non-commuta

Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention

Model ReleasesDGX agent

arXiv:2606.01243v1 Announce Type: new Abstract: Latent reasoning enables Large Language Models (LLMs) to perform multi-step inference within continuous hidden states, offering efficiency gains over ex

URDF-Anything+: End-to-End Generation for Simulation-Ready Articulated Assets

Model ReleasesDGX agent

arXiv:2603.14010v2 Announce Type: replace Abstract: Articulated objects are fundamental for robotics, simulation of physics, and interactive virtual environments. However, recovering them from visual

v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound

Model ReleasesDGX agent

arXiv:2509.25773v3 Announce Type: replace-cross Abstract: AI models capable of comprehending humor hold real-world promise -- for example, enhancing engagement in human-machine interactions. To gauge

v0.30.1: llm: ignore llama-server SSE ping comments (#16443)

Model ReleasesDGX agent

Ollama v0.30.1 addresses an issue where the LLM component now ignores Server-Sent Events (SSE) ping comments from llama-server, resolving problem #16443. This fix improves the stability and reliabilit

Value Flows

Model ReleasesDGX agent

arXiv:2510.07650v4 Announce Type: replace-cross Abstract: While most reinforcement learning methods today flatten the distribution of future returns to a single scalar value, distributional RL methods

VESTA: Visual Exploration with Statistical Tool Agents

Model ReleasesDGX agent

arXiv:2606.00384v1 Announce Type: new Abstract: Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated. Recent agent-based systems lev

Vision-language Models for Driver Monitoring Systems: A Driver Activity Description Dataset

Model ReleasesDGX agent

arXiv:2606.02273v1 Announce Type: new Abstract: Understanding subtle driver actions is essential for building reliable driver monitoring systems. Existing visionlanguage models (VLMs) are trained on g

Visual-Noise Guided In-Context Distillation for Multimodal Large Language Model Unlearning

Model ReleasesDGX agent

arXiv:2606.00105v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress on vision-language tasks, but they may also memorize and expose sensitive o

VLBM: Variational Latent Basis Modeling for OOD Robust Multivariate Time Series Forecasting

Model ReleasesDGX agent

arXiv:2606.02138v1 Announce Type: cross Abstract: Out of distribution (OOD) events in multivariate time series forecasting are rare but often dominate real world risk, making average case forecasting

VocSim: A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio

Model ReleasesDGX agent

arXiv:2512.10120v2 Announce Type: replace-cross Abstract: General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identit

WAON: A Large-Scale Japanese Image-Text Dataset for Cultural Adaptation in Contrastive Vision-Language Models

Model ReleasesDGX agent

arXiv:2510.22276v3 Announce Type: replace-cross Abstract: Contrastive vision-language models have achieved remarkable progress through large-scale pretraining. Recent work has shown that removing Engl

We wrapped a live session on M3 yesterday with the @togethercompute team & our researchers @zpysky1125 and @HaohaiSun A few highlights 🧵 1.…

Model ReleasesDGX agent

We wrapped a live session on M3 yesterday with the @togethercompute team & our researchers @zpysky1125 and @HaohaiSun A few highlights 🧵 1. MSA (MiniMax Sparse Attention) is the star ⭐️. Unlike CSA/HC

We're sponsoring a hackathon to scale down. Hosted by our friends @huggingface and @Gradio, we want working with models to feel like yours a…

Model ReleasesDGX agent

We're sponsoring a hackathon to scale down. Hosted by our friends @huggingface and @Gradio, we want working with models to feel like yours again. Small enough that it's inexpensive to run, big enough

What Do LLMs Know About Alzheimer's Disease? Multi-loss Fine-Tuning and Probing for AD Detection

Model ReleasesDGX agent

arXiv:2602.11177v2 Announce Type: replace-cross Abstract: Reliable early detection of Alzheimer's disease (AD) is challenging, particularly due to the limited availability of labeled data. While large

What to Format and How: A Benchmark and Workflow Approach for Document Formatting

Model ReleasesDGX agent

arXiv:2606.01936v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have opened up new possibilities for automated document formatting. However, real-world formatting often

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Model ReleasesDGX agent

arXiv:2602.16763v2 Announce Type: replace Abstract: Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks qui

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning

Model ReleasesDGX agent

arXiv:2602.08236v2 Announce Type: replace-cross Abstract: Despite rapid progress in MLLMs, visual spatial reasoning remains unreliable when correct answers depend on how a scene would appear under uns

When Jokes Cross the Line: Analyzing Regular Humor and Dark Humor in YouTube Shorts

Model ReleasesDGX agent

arXiv:2606.00046v1 Announce Type: cross Abstract: Video platforms such as YouTube have reshaped how users engage with entertainment and information, emphasizing brief, highly engaging content such as

← Previous
1…186187188189190…378
Next →