AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,595 results
22 May 2026

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

Model ReleasesDGX agent

arXiv:2605.20630v1 Announce Type: new Abstract: Industrial asset operations workflows are latency-sensitive because a single user query may require coordination over sensor data, work orders, failure

Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset

Model ReleasesDGX agent

arXiv:2605.22186v1 Announce Type: new Abstract: Event-based low-light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking t

EventGait: Towards Robust Gait Recognition with Event Streams

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.22139v1 Announce Type: new Abstract: Gait recognition enables non-intrusive, privacy-preserving identification but suffers in uncontrolled environments due to illumination and motion sensit

FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

Model ReleasesDGX agent

arXiv:2605.22552v1 Announce Type: new Abstract: Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

Model ReleasesDGX agent

arXiv:2605.21625v1 Announce Type: cross Abstract: The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus

FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments

Model ReleasesDGX agent

arXiv:2605.22018v1 Announce Type: new Abstract: The Flooded Road Environments Dataset (FRED) is, to our knowledge, the first multi-modal autonomous driving dataset specifically targeting the collectio

From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment

Model ReleasesDGX agent

arXiv:2605.21558v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains typically incurs high data and computational overhead. While prior efficiency efforts hav

From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding

Model ReleasesDGX agent

arXiv:2605.22413v1 Announce Type: new Abstract: Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation. While recent Multi

Gemini Flash 3.5 is such a disappointing model. It's intelligence and speed is awesome. Absolutely amazing. But it's been trained to max eva…

Model ReleasesDGX agent

Gemini Flash 3.5 is such a disappointing model. It's intelligence and speed is awesome. Absolutely amazing. But it's been trained to max evals, not to be helpful to humans. It goes off and does random

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning

Model ReleasesDGX agent

arXiv:2605.22558v1 Announce Type: new Abstract: Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearan

GHI: Graphormer over Conditioned Hypergraph Incidence for Aspect-Based Sentiment Analysis

Model ReleasesDGX agent

arXiv:2605.22228v1 Announce Type: new Abstract: Aspect-based sentiment analysis (ABSA) requires models to bind sentiment evidence to the correct aspect, making it a natural testbed for fine-grained st

GPT 5.5 seems to be improving in that direction now, and Claude models are getting worse at it, so I don't think there's a clear winner now.

Model ReleasesDGX agent

Jeremy Howard comments on comparative performance trends between GPT 5.5 and Claude models, noting that GPT 5.5 appears to be improving in a particular capability while Claude models are declining in

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

Model ReleasesDGX agent

arXiv:2605.20203v1 Announce Type: cross Abstract: As older adults increasingly use LLM-based chatbots for companionship and assistance, a safety gap is emerging. Older adults may face vulnerabilities

H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning

Model ReleasesDGX agent

arXiv:2605.22629v1 Announce Type: new Abstract: Parametric human models capture global pose but cannot represent the non-rigid surface dynamics of clothing and soft tissue. Generic scene flow estimate

ha ha, so much for step change? maybe this problem was just easier than some?

Model ReleasesDGX agent

ha ha, so much for step change? maybe this problem was just easier than some? The standard GPT-5.5 reproduced the proof ~ 👇 https://chatgpt.com/share/6a0e9e04-8cb0-8332-a4f1-ec68acd2e03e You don't nee

Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer

Model ReleasesDGX agent

arXiv:2605.22007v1 Announce Type: new Abstract: Hallucination is often viewed as a direct consequence of missing knowledge: a model answers incorrectly when the correct answer is absent from its gener

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

Model ReleasesDGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

Highlighting the new WebGPU backend in llama.cpp/ggml The work to bring full-fledged WebGPU support in llama.cpp started about an year and a…

Model ReleasesDGX agent

Highlighting the new WebGPU backend in llama.cpp/ggml The work to bring full-fledged WebGPU support in llama.cpp started about an year and a half ago. It has been lead by @reeselevine and team at USCS

Highlights from today’s Codex Thursday launches: 1️⃣ Codex can now securely use apps on your Mac from your phone, even when your Mac is lock…

Model ReleasesDGX agent

Highlights from today’s Codex Thursday launches: 1️⃣ Codex can now securely use apps on your Mac from your phone, even when your Mac is locked and the screen is off. http://developers.openai.com/codex

How Virgin Atlantic ships faster with Codex

Model ReleasesDGX agent

Virgin Atlantic implemented OpenAI's Codex to accelerate software development and deployment processes, enabling faster shipping of features and updates. The case study demonstrates how Codex improved

How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing

Model ReleasesDGX agent

arXiv:2602.01851v2 Announce Type: replace Abstract: Recent generative models have achieved remarkable progress in image editing. However, existing systems and benchmarks remain largely text-guided. In

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild

Model ReleasesDGX agent

arXiv:2605.22064v1 Announce Type: new Abstract: Hy-MT2 is a family of fast-thinking multilingual translation models designed for complex real-world scenarios. It includes three model sizes: 1.8B, 7B,

HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering

Model ReleasesDGX agent

arXiv:2605.22035v1 Announce Type: cross Abstract: Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge

I have to eat crow on this, in light of further information. whatever OpenAI spent on Erdos using a new model, apparently you can get GPT 5.…

Model ReleasesDGX agent

I have to eat crow on this, in light of further information. whatever OpenAI spent on Erdos using a new model, apparently you can get GPT 5.5 to do something similar; @emollick’s presumably estimates

I think people don't realize why Gemini Omni is different than other video AIs. It is fully multimodal, so it can edit video natively, too I…

Model ReleasesDGX agent

I think people don't realize why Gemini Omni is different than other video AIs. It is fully multimodal, so it can edit video natively, too I took the famous 'train ' movie from 1896 & made it a bullet

I truly miss the age of science and transparency in AI. Especially given how much money and political power and governance and scientific un…

Model ReleasesDGX agent

I truly miss the age of science and transparency in AI. Especially given how much money and political power and governance and scientific understanding is at stake. We don’t know for example • How man

IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions

Model ReleasesDGX agent

arXiv:2605.22247v1 Announce Type: new Abstract: Idioms pose a fundamental challenge for language models, as their meaning cannot be inferred from surface form alone. Understanding such expressions, th

InfVSR: Breaking Length Limits of Generic Video Super-Resolution

Model ReleasesDGX agent

arXiv:2510.00948v2 Announce Type: replace Abstract: Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent c

InnerQ: Hardware-Aware Tuning-Free Quantization of KV Cache for Large Language Models

Model ReleasesDGX agent

arXiv:2602.23200v2 Announce Type: replace-cross Abstract: When transformer-based language models are deployed for text generation, most of the inference time is spent in the decoding stage, where outp

InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation

Model ReleasesDGX agent

arXiv:2510.09724v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly capable of generating complete applications from natural language instructions, creating new opp

Introducing Qwen3.7-Max from @Alibaba_Qwen, Qwen’s flagship model for the agent era with 1M context and leading performance across agentic c…

Model ReleasesDGX agent

Introducing Qwen3.7-Max from @Alibaba_Qwen, Qwen’s flagship model for the agent era with 1M context and leading performance across agentic coding, reasoning, and long-horizon autonomy. AI natives can

Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements

Model ReleasesDGX agent

arXiv:2605.22079v1 Announce Type: new Abstract: Large language models (LLMs) are widely used to generate structured outputs such as JSON, SQL, and code, yet public resources remain limited for evaluat

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

Model ReleasesDGX agent

arXiv:2605.22080v1 Announce Type: new Abstract: We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF material

Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models

Model ReleasesDGX agent

arXiv:2605.21861v1 Announce Type: new Abstract: Multi-modality medical vision (MV) foundation models (FM) are fundamentally challenged by pronounced Non-IID feature statistics across heterogeneous ima

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.21988v1 Announce Type: new Abstract: Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

Model ReleasesDGX agent

arXiv:2605.21573v1 Announce Type: new Abstract: We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with

Linear Dynamics in the RLVR Training of Large Language Models

Model ReleasesDGX agent

arXiv:2601.04537v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven significant performance gains in reasoning-oriented large language models (LL

Link to GPT 5.5 on the recent Erdo problem: https://x.com/maxiao54704/status/2057484153755480537?s=61

Model ReleasesDGX agent

Link to GPT 5.5 on the recent Erdo problem: https://x.com/maxiao54704/status/2057484153755480537?s=61 The standard GPT-5.5 reproduced the proof ~ 👇 https://chatgpt.com/share/6a0e9e04-8cb0-8332-a4f1-ec

LLM Readiness Harness: Evaluation, Observability, and CI Gates for LLM/RAG Applications

Model ReleasesDGX agent

arXiv:2603.27355v2 Announce Type: replace-cross Abstract: We present a readiness harness for LLM and RAG applications that turns evaluation into a deployment decision workflow. The system combines aut

LongVT: Incentivizing 'Thinking with Long Videos' via Native Tool Calling

Model ReleasesDGX agent

arXiv:2511.20785v3 Announce Type: replace Abstract: Large multimodal models (LMMs) have shown great potential for video reasoning with textual Chain-of-Thought. However, they remain vulnerable to hall

LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model

Model ReleasesDGX agent

arXiv:2605.22089v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sp

M3: Conversational LLMs Simplify Secure Clinical Data Access, Understanding, and Analysis

Model ReleasesDGX agent

arXiv:2507.01053v4 Announce Type: replace-cross Abstract: Large-scale clinical databases offer opportunities for medical research, but their complexity creates barriers to effective use. The Medical I

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

Model ReleasesDGX agent

arXiv:2605.22177v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and modular skills has endowed autonomous agents with increasingly powerful capabilities. Existing f

MAP4TS: A Multi-Aspect Prompting Framework for Time-Series Forecasting with Large Language Models

Model ReleasesDGX agent

arXiv:2510.23090v2 Announce Type: replace Abstract: Recent advances have investigated the use of pretrained large language models (LLMs) for time-series forecasting by aligning numerical inputs with L

MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation

Model ReleasesDGX agent

arXiv:2605.22469v1 Announce Type: new Abstract: Evaluating single-concept personalization in text-to-image diffusion requires measuring both concept preservation, which captures identity fidelity to a

Matching with Deliberation: Test-Time Evolutionary Hierarchical Multi-Agents for Zero-Shot Compositional Image Retrieval

Model ReleasesDGX agent

arXiv:2605.22478v1 Announce Type: new Abstract: Zero-Shot Compositional Image Retrieval (ZS-CIR) requires both preserving the visual continuity of the reference image and faithfully executing the sema

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

Model ReleasesDGX agent

arXiv:2605.21917v1 Announce Type: new Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues

Model ReleasesDGX agent

arXiv:2605.21954v1 Announce Type: new Abstract: Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal la

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

Model ReleasesDGX agent

arXiv:2605.21796v1 Announce Type: cross Abstract: Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

Model ReleasesDGX agent

arXiv:2605.22818v1 Announce Type: new Abstract: Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally inco

MOTOR: A Multimodal Dataset for Two-Wheeler Rider Behavior Understanding

Model ReleasesDGX agent

arXiv:2605.22550v1 Announce Type: new Abstract: Two-wheelers account for a disproportionately high share of road fatalities in the Global South. Research on two-wheeler rider behavior, however, lags f

MTP means Multi Token Prediction. It's a speculative decoding technique that can result in large inference speedups in many cases. 1. Update…

Model ReleasesDGX agent

MTP means Multi Token Prediction. It's a speculative decoding technique that can result in large inference speedups in many cases. 1. Update to LM Studio 0.4.14 2. Download a model that supports MTP l

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

Model ReleasesDGX agent

arXiv:2505.17123v3 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly

My concern for the AI era, or at least this phase of it, is that a generation is being taught that 'close enough' is just fine. Take @Anthro…

Model ReleasesDGX agent

My concern for the AI era, or at least this phase of it, is that a generation is being taught that 'close enough' is just fine. Take @AnthropicAI for example. Text wrapping in Claude Code has been bro

Not All Starting Points Are Equal: Pre-trained Priors and Their Outsized Impact on Person Identification

Model ReleasesDGX agent

arXiv:2507.17640v3 Announce Type: replace Abstract: Recent years have seen an explosion of diverse general purpose pre-training methodologies for computer vision. However, the impact that these pre-tr

NuExtract3 released: open-weight 4B VLM for Markdown, OCR and structured extraction (self-hostable) [P]

Model ReleasesDGX agent

NuExtract3 is a unified 4B vision-language reasoning model for document understanding that combines structured information extraction with image-to-Markdown conversion, suitable for OCR and RAG prepro

one of the best decisions i made was to start tracing my claude code sessions into LangSmith. Has been a game changer to be able to share my…

Model ReleasesDGX agent

one of the best decisions i made was to start tracing my claude code sessions into LangSmith. Has been a game changer to be able to share my conversations, track usage patterns, monitor cost. And w ne

One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2605.22144v1 Announce Type: new Abstract: Existing approaches for digital short-drama production typically rely on one-shot LLM generated scripts and loosely coupled pipelines, which fail to sat

oops! wild update, strongly supports @emollick’s overall take:

Model ReleasesDGX agent

oops! wild update, strongly supports @emollick’s overall take: I have to eat crow on this, in light of further information. whatever OpenAI spent on Erdos using a new model, apparently you can get GPT

Open-World Evaluations for Measuring Frontier AI Capabilities

Model ReleasesDGX agent

arXiv:2605.20520v1 Announce Type: new Abstract: Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it

← Previous
1…220221222223224…377
Next →