AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,553 results
Model Releases

GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries

DGX agent

arXiv:2607.20757v1 Announce Type: cross Abstract: Transformers are known to have internal continuous symmetries that leave outputs invariant, while modifying quantization. GaugeQuant leverages this in

model-releasesarxiv-cs-cl
24 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Geometric Configurations of Perturbed Jailbreak Prompts

DGX agent

arXiv:2607.20581v1 Announce Type: cross Abstract: Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Get started with OpenAI GPT-5.6 Sol, Terra, and Luna on Amazon Bedrock

DGX agent

OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock. Learn how to select a model, run inference through the Responses API on the bedrock-mantle endpoint, reduce cost with

model-releasesaws-ml-blog
24 Jul 2026
Model Releases

Getting the most out of MTP

DGX agent

If you want to get the most out of MTP. You have to run some tests / benchmarks to do so. Turning it on with defaults will get improvements, but for many models and card combinations, you are leaving

model-releasesr-localllama
24 Jul 2026
Model Releases

GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis

DGX agent

arXiv:2607.18218v2 Announce Type: replace-cross Abstract: Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis, and

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Glad to see Google sharing data on how Gemini is being used. Especially interesting is that the usefulness of multimodal AI for manual labor…

DGX agent

Glad to see Google sharing data on how Gemini is being used. Especially interesting is that the usefulness of multimodal AI for manual labor may be greater than expected. https://blog.google/innovatio

model-releasesethan-mollick--x
24 Jul 2026
Model Releases

GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus

DGX agent

arXiv:2607.20443v1 Announce Type: new Abstract: We release GLAN-QnA-KR, a 303,581-row openly redistributable Korean instruction-QA corpus produced via the seedless taxonomy-driven GLAN synthesis pipel

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning

DGX agent

arXiv:2607.20730v1 Announce Type: cross Abstract: Large language models increasingly use search tools to retrieve up-to-date information, introducing a new attack surface in which retrieved documents

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Gradient Concentration, Not Weight Saliency, Explains Representation-Level Class Unlearning

DGX agent

arXiv:2607.21353v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of specific training data while preserving model utility. Many state-of-the-art approaches pursue this g

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis

DGX agent

arXiv:2607.21448v1 Announce Type: new Abstract: Dynamic scene reconstruction with 3D Gaussian Splatting requires a balance between fine-grained motion modeling, structural stability, and compact repre

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

GuardianAgentBench: Where Agents Fail and How to Guard Them

DGX agent

arXiv:2607.20982v1 Announce Type: new Abstract: As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavi

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

HalluScope: Fine-grained Hallucination Diagnosis for Multimodal Large Language Models

DGX agent

arXiv:2607.21105v1 Announce Type: new Abstract: Although Multimodal Large Language Models have achieved strong performance across a wide range of vision-language tasks, they still suffer from hallucin

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

Honest take on Laguna S2.1 and its uses (from actual use)

DGX agent

So I've taken some time to actually test laguna on a few of my own projects. I wanted to share as I feel most peoples comments at this point have just been about getting it running or saying it doesnt

model-releasesr-localllama
24 Jul 2026
Model Releases

How Many Bits Can an Adapter Write? Measuring the Capacity and Memorization of Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2607.21351v1 Announce Type: new Abstract: A LoRA adapter is a few megabytes that almost everyone treats as a skill rather than a record of the data behind it. We put that assumption on a scale.

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

HyperImageNet: A Large-Scale High-Spatial Resolution Hyperspectral Imagery Classification Benchmark

DGX agent

arXiv:2607.21050v1 Announce Type: new Abstract: We present HyperImageNet, a large-scale benchmark for fine-grained hyperspectral land-cover understanding. The dataset contains 26,084 airborne hyperspe

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

HypNO: A Graph-Based Neural Operator with Physics-Informed Message Passing for Hyperbolic Conservation Laws

DGX agent

arXiv:2607.20541v1 Announce Type: cross Abstract: We introduce HypNO, a graph-based neural operator for scalar hyperbolic conservation laws. HypNO operates directly on a space-time graph of finite-vol

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving

DGX agent

arXiv:2607.20988v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While pixel-level

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

I built a compiler that turns computation graphs into the weights of a vanilla transformer — no training anywhere [P]

DGX agent

I've been chasing the question of what algorithms a transformer can actually express -- separate from what it can learn. So I built a compiler: define a computation graph in ordinary Python, and it pr

model-releasesr-machinelearning
24 Jul 2026
Model Releases

I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]

DGX agent

Built an open-source AI coding agent that was 7%–75% cheaper than a cold 'claude -p' run on 6/6 well-localized tasks across repositories up to ~82k LOC. The biggest difference: Cold agent: 6.83, 207 t

model-releasesr-machinelearning
24 Jul 2026
Model Releases

ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders

DGX agent

arXiv:2607.21217v1 Announce Type: new Abstract: The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues

DGX agent

arXiv:2604.01925v2 Announce Type: replace-cross Abstract: Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit bias

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

In their letter defending open-weight AI, tech companies urge lawmakers to avoid 'premature restrictions' that 'stifle competition or drive innovation overseas' (Ashley Capoot/CNBC)

DGX agent

Ashley Capoot / CNBC: In their letter defending open-weight AI, tech companies urge lawmakers to avoid “premature restrictions” that “stifle competition or drive innovation overseas” — Nvidia, Microso

model-releasestechmeme
24 Jul 2026
Model Releases

Incomplete Prompt Jailbreaks in Large Language Models

DGX agent

arXiv:2607.20473v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, sentence completion

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

DGX agent

arXiv:2607.20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or n

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?

DGX agent

arXiv:2607.20460v1 Announce Type: cross Abstract: Current full-duplex (FD) spoken dialogue systems can produce fluid interactions, yet it remains unclear whether they can adapt their turn-taking behav

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Introducing Claude Opus 5

DGX agent

Introducing Claude Opus 5 I've been offline kayaking with sea otters for much of today so I haven't had a chance to put Anthropic's new model Claude Opus 5 through its paces yet. The buzz is positive,

model-releasessimon-willison
24 Jul 2026
Model Releases

Introducing Claude Opus 5 on AWS: Anthropic’s most capable Opus model

DGX agent

This post covers Opus 5’s improvements and practical guidance for AI engineers integrating the model into agentic systems and production inference workloads on Amazon Bedrock. See the documentation fo

model-releasesaws-ml-blog
24 Jul 2026
Model Releases

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

DGX agent

arXiv:2607.20891v1 Announce Type: new Abstract: Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence synthesis, and report generation, y

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought

DGX agent

arXiv:2607.20427v1 Announce Type: cross Abstract: Mixture-of-Experts architectures have revolutionized scaling, yet the underlying logic of their routing remains a black box. In this paper, we uncover

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation

DGX agent

arXiv:2607.20494v1 Announce Type: new Abstract: Production LLM applications commonly stack a regex filter in front of model-side alignment; prior work found no measurable coverage gain from adding a l

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

DGX agent

arXiv:2607.20759v1 Announce Type: cross Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with au

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

DGX agent

arXiv:2607.20466v1 Announce Type: new Abstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equiv

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Just over 6 months later, Opus 5 now produces near-superhuman level spreadsheets and slide decks that match what a consultant would make. Th…

DGX agent

Just over 6 months later, Opus 5 now produces near-superhuman level spreadsheets and slide decks that match what a consultant would make. Things are changing fast. Media I'm hearing from many folks ac

model-releasesboris-cherny--x
24 Jul 2026
Model Releases

.@Kimi_Moonshot K3 lands on Together on Monday! We ran 452 DeepSWE rollouts against Claude Fable 5: near-flagship coding at ~35% of the pric…

DGX agent

.@Kimi_Moonshot K3 lands on Together on Monday! We ran 452 DeepSWE rollouts against Claude Fable 5: near-flagship coding at ~35% of the price, and K3 pulls ahead at higher pass@k's. Full deep-dive: ht

model-releasestogether-ai--x
24 Jul 2026
Model Releases

Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing

DGX agent

arXiv:2607.20723v1 Announce Type: cross Abstract: This work presents LeakyLMs, a set of attacks that leak proprietary model, architecture, and deployment information from production language models. L

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc

DGX agent

arXiv:2607.20456v1 Announce Type: cross Abstract: Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such as MiniZinc

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Learning causality from internet videos in latent space first, and then using RL to teach the foundation model how to act. This approach is …

DGX agent

Learning causality from internet videos in latent space first, and then using RL to teach the foundation model how to act. This approach is 30× cheaper than Gemini 3.1 Flash on pretraining and achieve

model-releasesyann-lecun--x
24 Jul 2026
Model Releases

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters

DGX agent

arXiv:2607.11029v2 Announce Type: replace-cross Abstract: Recent progress in visual navigation has largely been driven by scale: end-to-end policies with hundreds of millions of parameters trained on

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

DGX agent

arXiv:2607.20872v1 Announce Type: new Abstract: Long-form legal research reports increasingly rely on LLMs and agentic research systems, but their reliability depends not only on answering the task, b

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time

DGX agent

arXiv:2603.20509v2 Announce Type: replace Abstract: Camera traps are vital for large-scale biodiversity monitoring, yet accurate automated analysis remains challenging due to diverse deployment enviro

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer

DGX agent

arXiv:2607.20924v1 Announce Type: new Abstract: Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion

model-releasesarxiv-cs-cv
24 Jul 2026
Model Releases

Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy

DGX agent

arXiv:2607.21372v1 Announce Type: cross Abstract: Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. While positivity guarantees

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education

DGX agent

arXiv:2607.21570v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise for medical education, but most existing systems focus on localized interactions such as question answering or

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Memory-Computation Tradeoffs in Semi Amortized Parametric Optimization

DGX agent

arXiv:2607.20769v1 Announce Type: new Abstract: Learning-enabled decision systems often use offline data or computation to reduce online compute cost. Despite the empirical success of such approaches,

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

MemTools: A Unified Research Framework for Interoperable Agent Memory

DGX agent

arXiv:2607.21404v1 Announce Type: new Abstract: While memory systems are essential for agent architectures, pervasive architectural fragmentation restricts systematic research. Existing implementation

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

Meta is making its AI chatbot more like an assistant

DGX agent

Meta is upgrading its AI chatbot with new productivity features in a bid to compete with rivals like Gemini, ChatGPT, and Claude. The update will allow Meta AI to tap into your calendar to help you pl

model-releasesthe-verge-ai
24 Jul 2026
Model Releases

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems

DGX agent

arXiv:2509.22047v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has been shown to be an effective algorithm when an accurate reward model is available. However, such a hi

model-releasesarxiv-cs-lg
24 Jul 2026
Model Releases

Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing

DGX agent

arXiv:2607.20433v1 Announce Type: cross Abstract: While language models remain frozen at their training state, the world evolves continuously. Knowledge editing has emerged as a key alternative to ful

model-releasesarxiv-cs-ai
24 Jul 2026
← Previous
1…8990919293…470
Next →