AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,340 results
Model Releases

GRASPrune: Global Gating for Budgeted Structured Pruning of Large Language Models

DGX agent

arXiv:2604.19398v1 Announce Type: new Abstract: Large language models (LLMs) are expensive to serve because model parameters, attention computation, and KV caches impose substantial memory and latency

model-releasesarxiv-cs-ai
22 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Guys, I am absolutely astounded. The Qwen 3.6 27b is like a jump to Qwen 4 from Qwen 27B 3.5. I just did a full suite of front end design te…

DGX agent

Guys, I am absolutely astounded. The Qwen 3.6 27b is like a jump to Qwen 4 from Qwen 27B 3.5. I just did a full suite of front end design tests and agentic benchmarks, made entirely by it. VERDICT: Th

model-releasesclem-delangue--x
22 Apr 2026
Model Releases

HalluAudio: A Comprehensive Benchmark for Hallucination Detection in Large Audio-Language Models

DGX agent

arXiv:2604.19300v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have recently achieved strong performance across various audio-centric tasks. However, hallucination, where models

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

HarDBench: A Benchmark for Draft-Based Co-Authoring Jailbreak Attacks for Safe Human-LLM Collaborative Writing

DGX agent

arXiv:2604.19274v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as co-authors in collaborative writing, where users begin with rough drafts and rely on LLMs to compl

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Harmful Intent as a Geometrically Recoverable Feature of LLM Residual Streams

DGX agent

arXiv:2604.18901v1 Announce Type: cross Abstract: Harmful intent is geometrically recoverable from large language model residual streams: as a linear direction in most layers, and as angular deviation

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

HarmoniDiff-RS: Training-Free Diffusion Harmonization for Satellite Image Composition

DGX agent

arXiv:2604.19392v1 Announce Type: new Abstract: Satellite image composition plays a critical role in remote sensing applications such as data augmentation, disaste simulation, and urban planning. We p

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

Has anyone seen ANY official communication from Anthropic or an Anthropic staff member about the fact that the checkbox for Claude Code on P…

DGX agent

Has anyone seen ANY official communication from Anthropic or an Anthropic staff member about the fact that the checkbox for Claude Code on Pro is back to being checked again, or is the only evidence t

model-releasessimon-willison--x
22 Apr 2026
Model Releases

Has Automated Essay Scoring Reached Sufficient Accuracy? Deriving Achievable QWK Ceilings from Classical Test Theory

DGX agent

arXiv:2604.19131v1 Announce Type: new Abstract: Automated essay scoring (AES) is commonly evaluated on public benchmarks using quadratic weighted kappa (QWK). However, because benchmark labels are ass

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

HELM: Harness-Enhanced Long-horizon Memory for Vision-Language-Action Manipulation

DGX agent

arXiv:2604.18791v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models fail systematically on long-horizon manipulation tasks despite strong short-horizon performance. We show that this

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Heterogeneity-Aware Personalized Federated Learning for Industrial Predictive Analytics

DGX agent

arXiv:2604.19451v1 Announce Type: new Abstract: Federated prognostics enable clients (e.g., companies, factories, and production lines) to collaboratively develop a failure time prediction model while

model-releasesarxiv-cs-lg
22 Apr 2026
Model Releases

How Adversarial Environments Mislead Agentic AI?

DGX agent

arXiv:2604.18874v1 Announce Type: new Abstract: Tool-integrated agents are deployed on the premise that external tools ground their outputs in reality. Yet this very reliance creates a critical attack

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

How Far Are Video Models from True Multimodal Reasoning?

DGX agent

arXiv:2604.19193v1 Announce Type: new Abstract: Despite remarkable progress toward general-purpose video models, a critical question remains unanswered: how far are these models from achieving true mu

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

How far are we from agents that can self-generate world knowledge? The work proposes an outcome-based reward that measures how much an agent…

DGX agent

How far are we from agents that can self-generate world knowledge? The work proposes an outcome-based reward that measures how much an agent's self-generated world knowledge actually improves its task

model-releasesdair-ai--x
22 Apr 2026
Model Releases

How to add an evaluation harness to your Gemini CLI coding agent

DGX agent

Coding agents can update prompts, wire in tools, and change application logic across your codebase in a single run. The hard part isn’t getting the agent to make changes, but... The post How to add an

model-releasesarize-ai
22 Apr 2026
Model Releases

HoWToBench: Holistic Evaluation for LLM's Capability in Human-level Writing using Tree of Writing

DGX agent

arXiv:2604.19071v1 Announce Type: new Abstract: Evaluating the writing capabilities of large language models (LLMs) remains a significant challenge due to the multidimensional nature of writing skills

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

HP-Edit: A Human-Preference Post-Training Framework for Image Editing

DGX agent

arXiv:2604.19406v1 Announce Type: cross Abstract: Common image editing tasks typically adopt powerful generative diffusion models as the leading paradigm for real-world content editing. Meanwhile, alt

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Hugging Face Releases ml-intern: An Open-Source AI Agent that Automates the LLM Post-Training Workflow [The 'AI Intern' that actually ships …

DGX agent

Hugging Face Releases ml-intern: An Open-Source AI Agent that Automates the LLM Post-Training Workflow [The 'AI Intern' that actually ships SOTA models ] This isn't just another ML Research Loop wrapp

model-releasesclem-delangue--x
22 Apr 2026
Model Releases

Human-Guided Harm Recovery for Computer Use Agents

DGX agent

arXiv:2604.18847v1 Announce Type: new Abstract: As LM agents gain the ability to execute actions on real computer systems, we need ways to not only prevent harmful actions at scale but also effectivel

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

I wonder if something happened to radicalize NoVa last year

DGX agent

I wonder if something happened to radicalize NoVa last year I can't believe Virginia dems got away with being so brazen. The 10-1 map looks absolutely insane and its wildly unfair and the public had f

model-releasesanthropic--x
22 Apr 2026
Model Releases

Image models tend to get much more stuck on a particular direction than text models, requiring clearing the context window fairly often. Per…

DGX agent

Image models tend to get much more stuck on a particular direction than text models, requiring clearing the context window fairly often. PerfectSquashBench is my new measure of how image models anchor

model-releasesethan-mollick--x
22 Apr 2026
Model Releases

Improving the Distributional Alignment of LLMs using Supervision

DGX agent

arXiv:2507.00439v4 Announce Type: replace Abstract: The ability to accurately align LLMs with diverse population groups on subjective questions would have great value. In this work, we show that addin

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

IndiaFinBench: An Evaluation Benchmark for Large Language Model Performance on Indian Financial Regulatory Text

DGX agent

arXiv:2604.19298v1 Announce Type: cross Abstract: We introduce IndiaFinBench, to our knowledge the first publicly available evaluation benchmark for assessing large language model (LLM) performance on

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

InsideOut: Measuring and Mitigating Insider-Outsider Bias in Interview Script Generation

DGX agent

arXiv:2509.21080v2 Announce Type: replace-cross Abstract: Advancements in Large language models (LLMs) have enabled a variety of downstream applications like story and interview script generation. How

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Intentional Updates for Streaming Reinforcement Learning

DGX agent

arXiv:2604.19033v1 Announce Type: cross Abstract: In gradient-based learning, a step size chosen in parameter units does not produce a predictable per-step change in function output. This often leads

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Introducing Gemini Enterprise Agent Platform, powering the next wave of agents

DGX agent

In the early days of generative AI, building safe and reliable business tools took massive engineering effort and a high tolerance for trial and error. We helped solve that with Vertex AI, our trusted

model-releasesgoogle-cloud-ai
22 Apr 2026
Model Releases

Introducing OpenAI Privacy Filter

DGX agent

OpenAI's Privacy Filter is a tool designed to help users protect sensitive information when using OpenAI's services by automatically detecting and redacting or filtering personally identifiable inform

model-releasesopenai
22 Apr 2026
Model Releases

Introducing the Google Cloud Knowledge Catalog

DGX agent

Traditional data catalogs were built as manual inventories for technical users, focusing on table structures rather than the deep context that AI agents need. When agents lack business semantics and d

model-releasesgoogle-cloud-ai
22 Apr 2026
Model Releases

Introducing workspace agents in ChatGPT

DGX agent

OpenAI introduced workspace agents in ChatGPT, a feature that enables AI assistants to perform tasks and take actions within connected workspace applications and tools. These agents can help automate

model-releasesopenai
22 Apr 2026
Model Releases

Introducing workspace agents in ChatGPT—shared agents that can handle complex tasks and long-running workflows across tools and teams.

DGX agent

OpenAI introduced workspace agents in ChatGPT, a new feature that enables shared agents capable of managing complex tasks and extended workflows across multiple tools and team members. These agents ar

model-releasesopenai--x
22 Apr 2026
Model Releases

Is Claude Code going to cost $100/month? Probably not - it's all very confusing

DGX agent

Anthropic today quietly (as in silently, no announcement anywhere at all) updated their claude.com/pricing page (but not their Choosing a Claude plan page, which shows up first for me on Google) to ad

model-releasessimon-willison
22 Apr 2026
Model Releases

Iterable launches Nova agent to assist markets in scaling customer personalization

DGX agent

Iterable Inc., a customer engagement platform, today announced that it launched an artificial intelligence agent designed to help marketers keep customer interactions relevant as campaigns scale. The

model-releasessiliconangle
22 Apr 2026
Model Releases

I've been mostly prompting the new model directly via the client.images.generate() API, I have no idea if that rewrites my prompts at all or…

DGX agent

I've been mostly prompting the new model directly via the client.images.generate() API, I have no idea if that rewrites my prompts at all or if it passes them straight to the model - details here http

model-releasessimon-willison--x
22 Apr 2026
Model Releases

just a friendly reminder that republicans have a trifecta and could pass a federal ban on gerrymandering *tomorrow* if they wanted to. but t…

DGX agent

just a friendly reminder that republicans have a trifecta and could pass a federal ban on gerrymandering *tomorrow* if they wanted to. but they don’t want to. they want to rig the system. and then the

model-releasesanthropic--x
22 Apr 2026
Model Releases

Kimi K2.6 is free on Nous Portal for the next 24 hours Made possible by @vercel's AI Gateway & @Kimi_Moonshot Run 'hermes update', then 'her…

DGX agent

Kimi K2.6 is free on Nous Portal for the next 24 hours Made possible by @vercel's AI Gateway & @Kimi_Moonshot Run 'hermes update', then 'hermes model' and select Kimi K2.6 to try out one of the most i

model-releaseskimi-moonshot--x
22 Apr 2026
Model Releases

Kimi K2.6 is now ranked #1 on OpenRouter's programming leaderboard.

DGX agent

Kimi K2.6, a model developed by Moonshot AI, has achieved the top ranking on OpenRouter's programming leaderboard, indicating superior performance in code generation and programming-related tasks comp

model-releaseskimi-moonshot--x
22 Apr 2026
Model Releases

LatticeVision: Image to Image Networks for Modeling Non-Stationary Spatial Data

DGX agent

arXiv:2505.09803v3 Announce Type: replace-cross Abstract: In many applications, we wish to fit a parametric statistical model to a small ensemble of spatially distributed random variables ('fields').

model-releasesarxiv-cs-lg
22 Apr 2026
Model Releases

LegalBench-BR: A Benchmark for Evaluating Large Language Models on Brazilian Legal Decision Classification

DGX agent

arXiv:2604.18878v1 Announce Type: new Abstract: We introduce LegalBench-BR, the first public benchmark for evaluating language models on Brazilian legal text classification. The dataset comprises 3,10

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

LePREC: Reasoning as Classification over Structured Factors for Assessing Relevance of Legal Issues

DGX agent

arXiv:2604.19464v1 Announce Type: cross Abstract: More than half of the global population struggles to meet their civil justice needs due to limited legal resources. While Large Language Models (LLMs)

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

Less Is More: Cognitive Load and the Single-Prompt Ceiling in LLM Mathematical Reasoning

DGX agent

arXiv:2604.18897v1 Announce Type: new Abstract: We present a systematic empirical study of prompt engineering for formal mathematical reasoning in the context of the SAIR Equational Theories Stage 1 c

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

Level Up Your Agents: Announcing Google's Official Skills Repository

DGX agent

As AI models improve, technical practitioners are increasingly turning to agentic AI tools to build with Google Cloud products, from Firebase and the Gemini API, to BigQuery and GKE. But how can you e

model-releasesgoogle-cloud-ai
22 Apr 2026
Model Releases

Lingua-SafetyBench: A Benchmark for Safety Evaluation of Multilingual Vision-Language Models

DGX agent

arXiv:2601.22737v2 Announce Type: replace Abstract: The robust safety of Vision-Language Large Models (VLLMs) against joint multilingual and multimodal threats remains severely underexplored. Current

model-releasesarxiv-cs-cv
22 Apr 2026
Model Releases

LiteParse, our OSS document parser, is really good at parsing complex PDF layouts, text, and tables into a clean spatial grid. The best part…

DGX agent

LiteParse, our OSS document parser, is really good at parsing complex PDF layouts, text, and tables into a clean spatial grid. The best part is it doesn't use VLMs or any ML models at all. It's entire

model-releasesjerry-liu--x
22 Apr 2026
Model Releases

LiveVLN: Breaking the Stop-and-Go Loop in Vision-Language Navigation

DGX agent

arXiv:2604.19536v1 Announce Type: new Abstract: Recent navigation systems achieve strong benchmark results, yet real-world deployment often remains visibly stop-and-go. This bottleneck arises because

model-releasesarxiv-cs-ro
22 Apr 2026
Model Releases

llama-server -hf ggml-org/Qwen3.6-27B-GGUF --spec-default

DGX agent

This post likely demonstrates running Qwen2 3.6B or 27B model in GGUF format using llama-server with default specifications, showcasing inference capabilities of quantized open-source models. The comm

model-releasesclem-delangue--x
22 Apr 2026
Model Releases

LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models

DGX agent

arXiv:2604.18803v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed in settings where reliable visual grounding carries operational consequences, yet their behavi

model-releasesarxiv-cs-ai
22 Apr 2026
Model Releases

LM Performance:With only 27B parameters, Qwen3.6-27B outperforms the Qwen3.5-397B-A17B (397B total / 17B active, ~15x larger!) on every majo…

DGX agent

LM Performance:With only 27B parameters, Qwen3.6-27B outperforms the Qwen3.5-397B-A17B (397B total / 17B active, ~15x larger!) on every major coding benchmark — including SWE-bench Verified (77.2 vs.

model-releasesqwen--x
22 Apr 2026
Model Releases

Lost in Translation: Do LVLM Judges Generalize Across Languages?

DGX agent

arXiv:2604.19405v1 Announce Type: new Abstract: Automatic evaluators such as reward models play a central role in the alignment and evaluation of large vision-language models (LVLMs). Despite their gr

model-releasesarxiv-cs-cl
22 Apr 2026
Model Releases

LoViF 2026 Challenge on Real-World All-in-One Image Restoration: Methods and Results

DGX agent

arXiv:2604.19445v1 Announce Type: new Abstract: This paper presents a review for the LoViF Challenge on Real-World All-in-One Image Restoration. The challenge aimed to advance research on real-world a

model-releasesarxiv-cs-cv
22 Apr 2026
← Previous
1…402403404405406…466
Next →