AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,612 results
Model Releases

Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems

DGX agent

arXiv:2605.22001v1 Announce Type: cross Abstract: Injection detectors deployed to protect LLM agents are calibrated on static, template-based payloads that announce themselves as override directives.

model-releasesarxiv-cs-cl
22 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety

DGX agent

arXiv:2605.22643v1 Announce Type: new Abstract: Background. Traditional safety benchmarks for language models evaluate generated text: whether a model outputs toxic language, reproduces bias, or follo

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Catch up on the Dialogues stage at Google I/O 2026.

DGX agent

The Dialogues stage at Google I/O 2026 brought together Google leaders, scientific minds and creative visionaries to discuss technological breakthroughs. Featured discussions included AI agents and pr

model-releasesgoogle-ai
22 May 2026
Model Releases

Check Your LLM's Secret Dictionary! Five Lines of Code Reveal What Your LLM Learned (Including What It Shouldn't Have)

DGX agent

arXiv:2605.22005v1 Announce Type: cross Abstract: We show that singular value decomposition of the lm_head} weight matrix of a transformer-based large language model -- requiring only five lines of Py

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoning

DGX agent

arXiv:2605.22734v1 Announce Type: new Abstract: Biomedical knowledge graphs (KGs) treat disease associations as static facts, but temporal information is crucial for clinical reasoning, e.g., a sympto

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Closed and gated surveillance economy SaaS systems and models will get replaced by open source in two tiers: - The models themselves - The e…

DGX agent

Closed and gated surveillance economy SaaS systems and models will get replaced by open source in two tiers: - The models themselves - The everything-ultra-app harness If an American open source champ

model-releasesyann-lecun--x
22 May 2026
Model Releases

COCOTree: A Dataset and Benchmark for Open Tree-Structured Visual Decomposition

DGX agent

arXiv:2605.22068v1 Announce Type: new Abstract: We formalize and enable the task of open tree decomposition, which segments an image into hierarchical trees of visual components with unconstrained gra

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Code Researcher: Deep Research Agent for Large Systems Code and Commit History

DGX agent

arXiv:2506.11060v2 Announce Type: replace-cross Abstract: Large Language Model (LLM)-based coding agents have shown promising results on coding benchmarks, but their effectiveness on systems code rema

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Cohesion-6K: An Arabic Dataset for Analyzing Social Cohesion and Conflict in Online Discourse

DGX agent

arXiv:2605.22447v1 Announce Type: new Abstract: The study of online discourse has become central to understanding societal polarization. While much research has focused on detecting overt toxicity, th

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Comparing LLM and Fine-Tuned Model Performance on NVDRS Circumstance Extraction with Varying Prompt Complexity

DGX agent

arXiv:2605.21845v1 Announce Type: new Abstract: Suicide is a leading cause of death in the United States, and understanding the circumstances that precede it requires extracting structured information

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

CoRMA: Contrastive RMA for Contact-Rich Meta-Adaptation

DGX agent

arXiv:2605.22082v1 Announce Type: new Abstract: We present CoRMA(Contrastive Robotic Motor Adaptation), a context-based meta-adaptation framework that modifies RMA for force-dominant assembly. CoRMA r

model-releasesarxiv-cs-ro
22 May 2026
Model Releases

Cross-Lingual Consensus: Aligning Multilingual Cultural Knowledge via Multilingual Self-Consistency

DGX agent

arXiv:2605.22137v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate strong capabilities across various tasks, they exhibit significant performance discrepancies across la

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models

DGX agent

arXiv:2605.21854v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have rapidly converged on a small set of architectural patterns: discrete-token autoregression (e.g. OpenVLA) and co

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

CTFExplorer: Evaluating LLM Offensive Agents Through Multi-Target Web CTF Benchmarking

DGX agent

arXiv:2602.08023v3 Announce Type: replace-cross Abstract: Existing benchmarks for LLM-based offensive security agents use isolated, single-target setups with a known vulnerable service and fixed objec

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Cursor Composer 2.5's is 3–18x cheaper than Opus 4.7 in Claude Code (medium reasoning), and 5–32x cheaper than GPT-5.5 in Codex (medium) bas…

DGX agent

Cursor Composer 2.5's is 3–18x cheaper than Opus 4.7 in Claude Code (medium reasoning), and 5–32x cheaper than GPT-5.5 in Codex (medium) based on API pricing This low Cost per Task isn't just driven b

model-releaseselon-musk--x
22 May 2026
Model Releases

Declarative Data Services: Structured Agentic Discovery for Composing Data Systems

DGX agent

arXiv:2605.20690v1 Announce Type: new Abstract: Agentic discovery has shown that LLM-driven search can find novel algorithms, designs, and code under benchmark conditions. Translating the paradigm to

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

DeepSeek v4: the most expected open-source model ever released, and the quietest landing

DGX agent

After 15 months of incremental updates, leaks, and rumored leaks, DeepSeek released version 4. It arrived without the fanfare R1 and R1-preview commanded in early 2025. That quiet reception is the mos

model-releaseslambda-labs
22 May 2026
Model Releases

DeepSeek’s New AI Is A Game Changer

DGX agent

DeepSeek released two advanced AI models—R1 and V3—with R1 specializing in complex reasoning and V3 designed for large-scale language processing. Notably, DeepSeek developed these high-performing mode

model-releasestwo-minute-papers
22 May 2026
Model Releases

DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation

DGX agent

arXiv:2605.21482v1 Announce Type: new Abstract: Deep research, in which an agent searches the open web, collects evidence, and derives an answer through extended reasoning, is a prominent use case for

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning

DGX agent

arXiv:2509.20912v4 Announce Type: replace Abstract: Recent advances in multimodal language models (MLLMs) have made thinking with images a dominant paradigm for multimodal reasoning. However, existing

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Dissecting Embodied Abilities in Multimodal Language Models through Skill-level Evaluation and Diagnosis

DGX agent

arXiv:2510.08759v2 Announce Type: replace Abstract: Understanding the capability bottlenecks of embodied multimodal large language models (MLLMs) is crucial for improving embodied agents. However, exi

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Diverse Yet Consistent: Context-Guided Diffusion with Energy-Based Joint Refinement for Multi-Agent Motion Prediction

DGX agent

arXiv:2605.22017v1 Announce Type: new Abstract: Deepgenerative models havebecomeapromisingapproach for human motion prediction due to their ability to capture multimodal distributions and represent di

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Do Vision Models Encode Object-Level Semantic Relatedness? A Cognitive Psychology-Inspired Benchmark

DGX agent

arXiv:1709.03806v2 Announce Type: replace Abstract: Modern vision models have achieved strong object-recognition performance, yet it remains unclear whether their representations encode object-level s

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Does Slightly Mean Somewhat? Measuring Vague Intensity Words in LLM Numeric Actions

DGX agent

arXiv:2605.21827v1 Announce Type: new Abstract: Do language models preserve the ordinal meaning of intensity words when those words must produce numeric actions? I study a researcher-constructed scale

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Dual-Integrated Low-Latency Single-Lens Infrared Computational Imaging for Object Detection

DGX agent

arXiv:2605.21964v1 Announce Type: new Abstract: Computational imaging enables compact infrared systems, but deep-learning pipelines that combine image reconstruction and object detection often introdu

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

DGX agent

arXiv:2605.22138v1 Announce Type: cross Abstract: How should an agent decide when and how to plan? A dominant approach builds agents as reactive policies with adaptive computation (e.g., chain-of-thou

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Essential context on OpenAI’s Erdos result

DGX agent

Essential context on OpenAI’s Erdos result I truly miss the age of science and transparency in AI. Especially given how much money and political power and governance and scientific understanding is at

model-releasesgary-marcus--x
22 May 2026
Model Releases

Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark

DGX agent

arXiv:2503.17599v3 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated considerable potential in general practice. However, existing benchmarks and evaluation frameworks pr

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Evaluating Commercial AI Chatbots as News Intermediaries

DGX agent

arXiv:2605.22785v1 Announce Type: new Abstract: AI chatbots are rapidly shaping how people encounter the news, yet no prior study has systematically measured how accurately these systems, with their p

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines

DGX agent

arXiv:2605.20630v1 Announce Type: new Abstract: Industrial asset operations workflows are latency-sensitive because a single user query may require coordination over sensor data, work orders, failure

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

Event-Illumination Collaborative Low-light Image Enhancement with a High-resolution Real-world Dataset

DGX agent

arXiv:2605.22186v1 Announce Type: new Abstract: Event-based low-light image enhancement (LIE) methods mainly focus on incorporating high dynamic range (HDR) information from events while overlooking t

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

EventGait: Towards Robust Gait Recognition with Event Streams

DGX agent

arXiv:2605.22139v1 Announce Type: new Abstract: Gait recognition enables non-intrusive, privacy-preserving identification but suffers in uncontrolled environments due to illumination and motion sensit

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning

DGX agent

arXiv:2605.22552v1 Announce Type: new Abstract: Fashion image retrieval is a cornerstone of modern e-commerce systems. A unified framework that supports diverse query formats and search intentions is

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

DGX agent

arXiv:2605.21625v1 Announce Type: cross Abstract: The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

FRED: A Multi-Modal Autonomous Driving Dataset for Flooded Road Environments

DGX agent

arXiv:2605.22018v1 Announce Type: new Abstract: The Flooded Road Environments Dataset (FRED) is, to our knowledge, the first multi-modal autonomous driving dataset specifically targeting the collectio

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

From Parameters to Data: A Task-Parameter-Guided Fine-Tuning Pipeline for Efficient LLM Alignment

DGX agent

arXiv:2605.21558v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains typically incurs high data and computational overhead. While prior efficiency efforts hav

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding

DGX agent

arXiv:2605.22413v1 Announce Type: new Abstract: Extracting structured information from visual documents (Visual Information Extraction, VIE) is a cornerstone of business automation. While recent Multi

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

Gemini Flash 3.5 is such a disappointing model. It's intelligence and speed is awesome. Absolutely amazing. But it's been trained to max eva…

DGX agent

Gemini Flash 3.5 is such a disappointing model. It's intelligence and speed is awesome. Absolutely amazing. But it's been trained to max evals, not to be helpful to humans. It goes off and does random

model-releasesjeremy-howard--x
22 May 2026
Model Releases

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning

DGX agent

arXiv:2605.22558v1 Announce Type: new Abstract: Spatio-temporal reasoning in vision-language models requires visual representations that preserve physical geometry rather than merely semantic appearan

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

GHI: Graphormer over Conditioned Hypergraph Incidence for Aspect-Based Sentiment Analysis

DGX agent

arXiv:2605.22228v1 Announce Type: new Abstract: Aspect-based sentiment analysis (ABSA) requires models to bind sentiment evidence to the correct aspect, making it a natural testbed for fine-grained st

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

GPT 5.5 seems to be improving in that direction now, and Claude models are getting worse at it, so I don't think there's a clear winner now.

DGX agent

Jeremy Howard comments on comparative performance trends between GPT 5.5 and Claude models, noting that GPT 5.5 appears to be improving in a particular capability while Claude models are declining in

model-releasesjeremy-howard--x
22 May 2026
Model Releases

GrandGuard: Taxonomy, Benchmark, and Safeguards for Elderly-Chatbot Interaction Safety

DGX agent

arXiv:2605.20203v1 Announce Type: cross Abstract: As older adults increasingly use LLM-based chatbots for companionship and assistance, a safety gap is emerging. Older adults may face vulnerabilities

model-releasesarxiv-cs-ai
22 May 2026
Model Releases

H-Flow: Self-supervised Human Scene Flow via Physics-inspired Joint Multi-modal Learning

DGX agent

arXiv:2605.22629v1 Announce Type: new Abstract: Parametric human models capture global pose but cannot represent the non-rigid surface dynamics of clothing and soft tissue. Generic scene flow estimate

model-releasesarxiv-cs-cv
22 May 2026
Model Releases

ha ha, so much for step change? maybe this problem was just easier than some?

DGX agent

ha ha, so much for step change? maybe this problem was just easier than some? The standard GPT-5.5 reproduced the proof ~ 👇 https://chatgpt.com/share/6a0e9e04-8cb0-8332-a4f1-ec68acd2e03e You don't nee

model-releasesgary-marcus--x
22 May 2026
Model Releases

Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer

DGX agent

arXiv:2605.22007v1 Announce Type: new Abstract: Hallucination is often viewed as a direct consequence of missing knowledge: a model answers incorrectly when the correct answer is absent from its gener

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

HealthCraft: A Reinforcement Learning Safety Environment for Emergency Medicine

DGX agent

arXiv:2605.21496v1 Announce Type: cross Abstract: Frontier language models are being deployed into clinical workflows faster than the infrastructure to evaluate them safely. Static medical-QA benchmar

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Highlighting the new WebGPU backend in llama.cpp/ggml The work to bring full-fledged WebGPU support in llama.cpp started about an year and a…

DGX agent

Highlighting the new WebGPU backend in llama.cpp/ggml The work to bring full-fledged WebGPU support in llama.cpp started about an year and a half ago. It has been lead by @reeselevine and team at USCS

model-releasesgeorgi-gerganov--x
22 May 2026
Model Releases

Highlights from today’s Codex Thursday launches: 1️⃣ Codex can now securely use apps on your Mac from your phone, even when your Mac is lock…

DGX agent

Highlights from today’s Codex Thursday launches: 1️⃣ Codex can now securely use apps on your Mac from your phone, even when your Mac is locked and the screen is off. http://developers.openai.com/codex

model-releasesopenai--x
22 May 2026
← Previous
1…275276277278279…472
Next →