AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,770 results
Model Releases

FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments

DGX agent

arXiv:2605.22018v1 Announce Type: new Abstract: The Flooded Road Environments Dataset (FRED) is, to our knowledge, the first multi-modal autonomous driving dataset specifically targeting the collectio

model-releasesarxiv-cs-cv
22 May 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment

DGX agent

arXiv:2605.21558v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains typically incurs high data and computational overhead. While prior efficiency efforts hav

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding

DGX agent

arXiv:2605.22413v1 Announce Type: new Abstract: Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation. While recent Multi

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Gemini Flash 3.5 is such a disappointing model. It's intelligence and speed is awesome. Absolutely amazing. But it's been trained to max eva…

DGX agent

Gemini Flash 3.5 is such a disappointing model. It's intelligence and speed is awesome. Absolutely amazing. But it's been trained to max evals, not to be helpful to humans. It goes off and does random

model-releasesjeremy-howard--x
22 May 2026
Model Releases

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning

DGX agent

arXiv:2605.22558v1 Announce Type: new Abstract: Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearan

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

GHI: Graphormer over Conditioned Hypergraph Incidence for Aspect-Based Sentiment Analysis

DGX agent

arXiv:2605.22228v1 Announce Type: new Abstract: Aspect-based sentiment analysis (ABSA) requires models to bind sentiment evidence to the correct aspect, making it a natural testbed for fine-grained st

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

GPT 5.5 seems to be improving in that direction now, and Claude models are getting worse at it, so I don't think there's a clear winner now.

DGX agent

Jeremy Howard comments on comparative performance trends between GPT 5.5 and Claude models, noting that GPT 5.5 appears to be improving in a particular capability while Claude models are declining in

model-releasesjeremy-howard--x
22 May 2026
Model Releases

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

DGX agent

arXiv:2605.20203v1 Announce Type: cross Abstract: As older adults increasingly use LLM-based chatbots for companionship and assistance, a safety gap is emerging. Older adults may face vulnerabilities

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning

DGX agent

arXiv:2605.22629v1 Announce Type: new Abstract: Parametric human models capture global pose but cannot represent the non-rigid surface dynamics of clothing and soft tissue. Generic scene flow estimate

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

ha ha, so much for step change? maybe this problem was just easier than some?

DGX agent

ha ha, so much for step change? maybe this problem was just easier than some? The standard GPT-5.5 reproduced the proof ~ 👇 https://chatgpt.com/share/6a0e9e04-8cb0-8332-a4f1-ec68acd2e03e You don't nee

model-releasesgary-marcus--x
22 May 2026
Model Releases

Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer

DGX agent

arXiv:2605.22007v1 Announce Type: new Abstract: Hallucination is often viewed as a direct consequence of missing knowledge: a model answers incorrectly when the correct answer is absent from its gener

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

DGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Highlighting the new WebGPU backend in llama.cpp/ggml The work to bring full-fledged WebGPU support in llama.cpp started about an year and a…

DGX agent

Highlighting the new WebGPU backend in llama.cpp/ggml The work to bring full-fledged WebGPU support in llama.cpp started about an year and a half ago. It has been lead by @reeselevine and team at USCS

model-releasesgeorgi-gerganov--x
22 May 2026
Model Releases

Highlights from today’s Codex Thursday launches: 1️⃣ Codex can now securely use apps on your Mac from your phone, even when your Mac is lock…

DGX agent

Highlights from today’s Codex Thursday launches: 1️⃣ Codex can now securely use apps on your Mac from your phone, even when your Mac is locked and the screen is off. http://developers.openai.com/codex

model-releasesopenai--x
22 May 2026
Model Releases

How Virgin Atlantic ships faster with Codex

DGX agent

Virgin Atlantic implemented OpenAI's Codex to accelerate software development and deployment processes, enabling faster shipping of features and updates. The case study demonstrates how Codex improved

model-releasesopenai
22 May 2026
Model Releases

How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing

DGX agent

arXiv:2602.01851v2 Announce Type: replace Abstract: Recent generative models have achieved remarkable progress in image editing. However, existing systems and benchmarks remain largely text-guided. In

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild

DGX agent

arXiv:2605.22064v1 Announce Type: new Abstract: Hy-MT2 is a family of fast-thinking multilingual translation models designed for complex real-world scenarios. It includes three model sizes: 1.8B, 7B,

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering

DGX agent

arXiv:2605.22035v1 Announce Type: cross Abstract: Continual Visual Question Answering (VQA) requires learning from non-stationary streams of visual inputs and questions while preserving past knowledge

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

I have to eat crow on this, in light of further information. whatever OpenAI spent on Erdos using a new model, apparently you can get GPT 5.…

DGX agent

I have to eat crow on this, in light of further information. whatever OpenAI spent on Erdos using a new model, apparently you can get GPT 5.5 to do something similar; @emollick’s presumably estimates

model-releasesgary-marcus--x
22 May 2026
Model Releases

I think people don't realize why Gemini Omni is different than other video AIs. It is fully multimodal, so it can edit video natively, too I…

DGX agent

I think people don't realize why Gemini Omni is different than other video AIs. It is fully multimodal, so it can edit video natively, too I took the famous 'train ' movie from 1896 & made it a bullet

model-releasesethan-mollick--x
22 May 2026
Model Releases

I truly miss the age of science and transparency in AI. Especially given how much money and political power and governance and scientific un…

DGX agent

I truly miss the age of science and transparency in AI. Especially given how much money and political power and governance and scientific understanding is at stake. We don’t know for example • How man

model-releasesgary-marcus--x
22 May 2026
Model Releases

IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions

DGX agent

arXiv:2605.22247v1 Announce Type: new Abstract: Idioms pose a fundamental challenge for language models, as their meaning cannot be inferred from surface form alone. Understanding such expressions, th

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

InfVSR: Breaking Length Limits of Generic Video Super-Resolution

DGX agent

arXiv:2510.00948v2 Announce Type: replace Abstract: Real-world videos often extend over thousands of frames. Existing generative video super-resolution (VSR) approaches, however, face two persistent c

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

InnerQ: Hardware-Aware Tuning-Free Quantization of KV Cache for Large Language Models

DGX agent

arXiv:2602.23200v2 Announce Type: replace-cross Abstract: When transformer-based language models are deployed for text generation, most of the inference time is spent in the decoding stage, where outp

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation

DGX agent

arXiv:2510.09724v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly capable of generating complete applications from natural language instructions, creating new opp

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Introducing Qwen3.7-Max from @Alibaba_Qwen, Qwen’s flagship model for the agent era with 1M context and leading performance across agentic c…

DGX agent

Introducing Qwen3.7-Max from @Alibaba_Qwen, Qwen’s flagship model for the agent era with 1M context and leading performance across agentic coding, reasoning, and long-horizon autonomy. AI natives can

model-releasestogether-ai--x
22 May 2026
Model Releases

Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements

DGX agent

arXiv:2605.22079v1 Announce Type: new Abstract: Large language models (LLMs) are widely used to generate structured outputs such as JSON, SQL, and code, yet public resources remain limited for evaluat

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

DGX agent

arXiv:2605.22080v1 Announce Type: new Abstract: We introduce JMed48k, a multi-profession Japanese healthcare licensing benchmark for evaluating vision-language models. Built from official PDF material

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models

DGX agent

arXiv:2605.21861v1 Announce Type: new Abstract: Multi-modality medical vision (MV) foundation models (FM) are fundamentally challenged by pronounced Non-IID feature statistics across heterogeneous ima

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning

DGX agent

arXiv:2605.21988v1 Announce Type: new Abstract: Video large language models (Video LLMs) achieve strong benchmark accuracy, yet often answer video questions through shortcuts such as single-frame cues

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models

DGX agent

arXiv:2605.21573v1 Announce Type: new Abstract: We introduce Lens, a 3.8B-parameter T2I model that achieves performance competitive with, and in several cases surpassing, state-of-the-art models with

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Linear Dynamics in the RLVR Training of Large Language Models

DGX agent

arXiv:2601.04537v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven significant performance gains in reasoning-oriented large language models (LL

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Link to GPT 5.5 on the recent Erdo problem: https://x.com/maxiao54704/status/2057484153755480537?s=61

DGX agent

Link to GPT 5.5 on the recent Erdo problem: https://x.com/maxiao54704/status/2057484153755480537?s=61 The standard GPT-5.5 reproduced the proof ~ 👇 https://chatgpt.com/share/6a0e9e04-8cb0-8332-a4f1-ec

model-releasesgary-marcus--x
22 May 2026
Model Releases

LLM Readiness Harness: Evaluation, Observability, and CI Gates for LLM/RAG Applications

DGX agent

arXiv:2603.27355v2 Announce Type: replace-cross Abstract: We present a readiness harness for LLM and RAG applications that turns evaluation into a deployment decision workflow. The system combines aut

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

LongVT: Incentivizing 'Thinking with Long Videos' via Native Tool Calling

DGX agent

arXiv:2511.20785v3 Announce Type: replace Abstract: Large multimodal models (LMMs) have shown great potential for video reasoning with textual Chain-of-Thought. However, they remain vulnerable to hall

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model

DGX agent

arXiv:2605.22089v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a promising framework for end-to-end autonomous driving. However, existing VLAs typically rely on sp

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

M3: Conversational LLMs Simplify Secure Clinical Data Access, Understanding, and Analysis

DGX agent

arXiv:2507.01053v4 Announce Type: replace-cross Abstract: Large-scale clinical databases offer opportunities for medical research, but their complexity creates barriers to effective use. The Medical I

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

DGX agent

arXiv:2605.22177v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and modular skills has endowed autonomous agents with increasingly powerful capabilities. Existing f

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

MAP4TS: A Multi-Aspect Prompting Framework for Time-Series Forecasting with Large Language Models

DGX agent

arXiv:2510.23090v2 Announce Type: replace Abstract: Recent advances have investigated the use of pretrained large language models (LLMs) for time-series forecasting by aligning numerical inputs with L

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation

DGX agent

arXiv:2605.22469v1 Announce Type: new Abstract: Evaluating single-concept personalization in text-to-image diffusion requires measuring both concept preservation, which captures identity fidelity to a

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Matching with Deliberation: Test-Time Evolutionary Hierarchical Multi-Agents for Zero-Shot Compositional Image Retrieval

DGX agent

arXiv:2605.22478v1 Announce Type: new Abstract: Zero-Shot Compositional Image Retrieval (ZS-CIR) requires both preserving the visual continuity of the reference image and faithfully executing the sema

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks

DGX agent

arXiv:2605.21917v1 Announce Type: new Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues

DGX agent

arXiv:2605.21954v1 Announce Type: new Abstract: Video temporal grounding (VTG), which localizes the start and end times of a queried event in an untrimmed video, is a key test of whether multimodal la

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue

DGX agent

arXiv:2605.21796v1 Announce Type: cross Abstract: Grounding language in the physical world requires AI systems to interpret references that emerge dynamically during conversation. While current vision

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning

DGX agent

arXiv:2605.22818v1 Announce Type: new Abstract: Current motion-controlled image-to-video generation models rigidly follow user-provided trajectories that are often sparse, imprecise, and causally inco

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MOTOR: A Multimodal Dataset for Two-Wheeler Rider Behavior Understanding

DGX agent

arXiv:2605.22550v1 Announce Type: new Abstract: Two-wheelers account for a disproportionately high share of road fatalities in the Global South. Research on two-wheeler rider behavior, however, lags f

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

MTP means Multi Token Prediction. It's a speculative decoding technique that can result in large inference speedups in many cases. 1. Update…

DGX agent

MTP means Multi Token Prediction. It's a speculative decoding technique that can result in large inference speedups in many cases. 1. Update to LM Studio 0.4.14 2. Download a model that supports MTP l

model-releaseslm-studio--x
22 May 2026
Model Releases

MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation

DGX agent

arXiv:2505.17123v3 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have shown promising results in complex reasoning tasks. However, current evaluations predominantly

model-releasesarxiv-cs-cl
22 May 2026
← Previous
1…279280281282283…475
Next →