AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,512 results
24 Jul 2026

Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation

Model ReleasesDGX agent

arXiv:2607.20494v1 Announce Type: new Abstract: Production LLM applications commonly stack a regex filter in front of model-side alignment; prior work found no measurable coverage gain from adding a l

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

Model ReleasesDGX agent

arXiv:2607.20759v1 Announce Type: cross Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with au

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.20466v1 Announce Type: new Abstract: Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equiv

Just over 6 months later, Opus 5 now produces near-superhuman level spreadsheets and slide decks that match what a consultant would make. Th…

Model ReleasesDGX agent

Just over 6 months later, Opus 5 now produces near-superhuman level spreadsheets and slide decks that match what a consultant would make. Things are changing fast. Media I'm hearing from many folks ac

.@Kimi_Moonshot K3 lands on Together on Monday! We ran 452 DeepSWE rollouts against Claude Fable 5: near-flagship coding at ~35% of the pric…

Model ReleasesDGX agent

.@Kimi_Moonshot K3 lands on Together on Monday! We ran 452 DeepSWE rollouts against Claude Fable 5: near-flagship coding at ~35% of the price, and K3 pulls ahead at higher pass@k's. Full deep-dive: ht

Leaky Language Models: Stealing Architecture and Inference Optimizations via Per-Token Timing

Model ReleasesDGX agent

arXiv:2607.20723v1 Announce Type: cross Abstract: This work presents LeakyLMs, a set of attacks that leak proprietary model, architecture, and deployment information from production language models. L

Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc

Model ReleasesDGX agent

arXiv:2607.20456v1 Announce Type: cross Abstract: Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such as MiniZinc

Learning causality from internet videos in latent space first, and then using RL to teach the foundation model how to act. This approach is …

Model ReleasesDGX agent

Learning causality from internet videos in latent space first, and then using RL to teach the foundation model how to act. This approach is 30× cheaper than Gemini 3.1 Flash on pretraining and achieve

Learning to Navigate Efficiently with Only 0.58M Trainable Parameters

Model ReleasesDGX agent

arXiv:2607.11029v2 Announce Type: replace-cross Abstract: Recent progress in visual navigation has largely been driven by scale: end-to-end policies with hundreds of millions of parameters trained on

LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports

Model ReleasesDGX agent

arXiv:2607.20872v1 Announce Type: new Abstract: Long-form legal research reports increasingly rely on LLMs and agentic research systems, but their reliability depends not only on answering the task, b

Lessons and Open Questions from a Unified Study of Camera-Trap Species Recognition Over Time

Model ReleasesDGX agent

arXiv:2603.20509v2 Announce Type: replace Abstract: Camera traps are vital for large-scale biodiversity monitoring, yet accurate automated analysis remains challenging due to diverse deployment enviro

MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer

Model ReleasesDGX agent

arXiv:2607.20924v1 Announce Type: new Abstract: Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion

Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy

Model ReleasesDGX agent

arXiv:2607.21372v1 Announce Type: cross Abstract: Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. While positivity guarantees

MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education

Model ReleasesDGX agent

arXiv:2607.21570v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise for medical education, but most existing systems focus on localized interactions such as question answering or

Memory-Computation Tradeoffs in Semi Amortized Parametric Optimization

Model ReleasesDGX agent

arXiv:2607.20769v1 Announce Type: new Abstract: Learning-enabled decision systems often use offline data or computation to reduce online compute cost. Despite the empirical success of such approaches,

MemTools: A Unified Research Framework for Interoperable Agent Memory

Model ReleasesDGX agent

arXiv:2607.21404v1 Announce Type: new Abstract: While memory systems are essential for agent architectures, pervasive architectural fragmentation restricts systematic research. Existing implementation

Meta is making its AI chatbot more like an assistant

Model ReleasesDGX agent

Meta is upgrading its AI chatbot with new productivity features in a bid to compete with rivals like Gemini, ChatGPT, and Claude. The update will allow Meta AI to tap into your calendar to help you pl

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems

Model ReleasesDGX agent

arXiv:2509.22047v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has been shown to be an effective algorithm when an accurate reward model is available. However, such a hi

Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing

Model ReleasesDGX agent

arXiv:2607.20433v1 Announce Type: cross Abstract: While language models remain frozen at their training state, the world evolves continuously. Knowledge editing has emerged as a key alternative to ful

Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation

Model ReleasesDGX agent

arXiv:2607.20908v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for optimized code

Multimodal CoLRAG-TF: Triple-Filtered Retrieval for Complex PDFs

Model ReleasesDGX agent

arXiv:2607.20517v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) over heterogeneous PDF collections remains challenging due to multimodal content, domain-specific terminology, and

Multimodal Pretraining for Generalizable EEG Representation Learning

Model ReleasesDGX agent

arXiv:2607.21384v1 Announce Type: new Abstract: Electroencephalography (EEG) models used for epilepsy are often limited to specific datasets and tasks. This limited approach can make it challenging to

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement

Model ReleasesDGX agent

arXiv:2607.21061v1 Announce Type: new Abstract: Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward

Naver-News-KO: A Korean News Summarization Dataset for Open-Source Fine-Tuning of Summarization Models

Model ReleasesDGX agent

arXiv:2607.20442v1 Announce Type: new Abstract: We release Naver-News-KO, a Korean news summarization dataset of 27,400 (document, summary) pairs collected from Naver News over a ten-day window in Jul

New research with Microsoft's and colleagues on training agents inside the harnesses they actually run in. (bookmark it) Why it matters: Age…

Model ReleasesDGX agent

New research with Microsoft's and colleagues on training agents inside the harnesses they actually run in. (bookmark it) Why it matters: Agents today live inside elaborate harnesses like Claude Code,

Non-Stationary Functional Bilevel Optimization

Model ReleasesDGX agent

arXiv:2601.15363v2 Announce Type: replace-cross Abstract: Functional bilevel optimization (FBO) provides a powerful framework for hierarchical learning in function spaces, yet current methods are limi

Nvidia releases Qwen-Image-Flash

Model ReleasesDGX agent

'The NVIDIA Qwen-Image-Flash model generates images from text prompts using a four-step, DMD2-distilled version of Qwen/Qwen-Image. The distillation used DMD2 from NVIDIA FastGen, NVIDIA Model Optimiz

ODeform: Learning Continuous 4D Motion for Shape Deformation with Neural ODEs

Model ReleasesDGX agent

arXiv:2607.20670v1 Announce Type: new Abstract: Modeling continuous object deformation is important for many computer vision and robotics tasks, such as manipulation and simulation. Existing approache

On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and quality, Opus 5 scores 63.6% with a 69.6% p…

Model ReleasesDGX agent

On FrontierCode 1.1, our benchmark for real-world engineering tasks that grades mergeability and quality, Opus 5 scores 63.6% with a 69.6% pass rate on Extended — approaching Fable 5 at half the cost.

One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies

Model ReleasesDGX agent

arXiv:2607.21143v1 Announce Type: cross Abstract: Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask, what to a

Open Knowledge format v0.2 tackles agentic trust

Model ReleasesDGX agent

When we introduced the Open Knowledge Format (OKF) in June 2026, we asserted that the context that agents need (table schemas, metric definitions, runbooks) should live in a format, not in a proprieta

Open Source Tax Engine outperforming fable 5 and gpt sol

Model ReleasesDGX agent

This is an open source and free tax engine which scored 96% on TaxCalcBench [highest ever recorded score till date] surpassing fable 5 and sol with just sonnet 5. The only 2 cases where it missed, it

OpenForgeRL: Train Harness-native Agents in Any Environment

Model ReleasesDGX agent

arXiv:2607.21557v1 Announce Type: new Abstract: Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to e

Optimizing an Ollama (Qwen:2.5) AI Agent: Fixing Search Aggregation, Context Bleed, and Query Extraction

Model ReleasesDGX agent

I am building a domain-specific AI agent powered by Ollama (using the qwen:2.5 model). For data retrieval, the agent utilizes multiple search APIs: DuckDuckGo Search (DDGS), Tavily, Serper, and Google

Opus 5 improves coding, reasoning efficiency, and prompt-cache-friendly tool use, and is priced at 5/1M input tokens and 25/1M output, same as Opus 4.8 (David Gewirtz/ZDNET)

Model ReleasesDGX agent

David Gewirtz / ZDNET: Opus 5 improves coding, reasoning efficiency, and prompt-cache-friendly tool use, and is priced at 5/1M input tokens and 25/1M output, same as Opus 4.8 — ZDNET's key takeaways —

Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most excitin…

Model ReleasesDGX agent

Opus 5 is a great model for coding, data analysis, design, biology, knowledge work. More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt inject

Opus 5 now available in Hermes Agent

Model ReleasesDGX agent

Claude Opus 5 is now released in the Hermes Agent, a product of Nous Research and Teknium. Users can access the model through multiple gateways, including the Nous Portal, OpenRouter, and Anthropic Di

pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development

Model ReleasesDGX agent

arXiv:2607.21268v1 Announce Type: cross Abstract: In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable co

[Paper] Statistically-Lossless Quantization of Large Language Models

Model ReleasesDGX agent

Model quantization has become essential for efficient large language model deployment, yet existing approaches involve clear trade-offs: methods such as GPTQ and AWQ achieve practical compression but

People are using Minecraft farms as AI agent benchmarks

Model ReleasesDGX agent

Someone modelled sugarcane farming as an integer program. See, sugarcane only grows next to water. Water costs one tile and can feed at most four cane tiles. The layout therefore becomes a coverage pr

PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails

Model ReleasesDGX agent

arXiv:2607.20482v1 Announce Type: new Abstract: Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provide underspeci

PhantomFill: When the Form Demands an Answer, Language Models Invent One

Model ReleasesDGX agent

arXiv:2607.20492v1 Announce Type: cross Abstract: Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the form itself

PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses

Model ReleasesDGX agent

arXiv:2603.13026v2 Announce Type: replace Abstract: Prompt injection poses serious security risks to real-world LLM applications, particularly autonomous agents. Although many defenses have been propo

Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM Benchmarks

Model ReleasesDGX agent

arXiv:2607.20864v1 Announce Type: cross Abstract: Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single ans

Preference Tuning as Spectral Update Reorganization

Model ReleasesDGX agent

arXiv:2607.20438v1 Announce Type: cross Abstract: Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains largely opa

ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning

Model ReleasesDGX agent

arXiv:2607.21022v1 Announce Type: new Abstract: Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing t

Profiling Lightweight Large Language Models

Model ReleasesDGX agent

arXiv:2607.20806v1 Announce Type: new Abstract: Lightweight large language models (LLMs) are increasingly being deployed locally on personal computers and are expected to play a growing role in resour

PromptPack: Scaling LLM Annotation Agents for Online Recommendation

Model ReleasesDGX agent

arXiv:2607.20528v1 Announce Type: new Abstract: Online recommendation platforms increasingly use Large Language Models (LLMs) to extract structured features from ad creatives. While deploying a single

QuantiBias: Benchmarking Quantization-Induced Bias in LLMs

Model ReleasesDGX agent

arXiv:2607.21063v1 Announce Type: new Abstract: Almost every large language model that reaches a broad audience is quantized: trained in full precision, then compressed for efficiency. This step is as

RE-AD: Real-Time Requirement Adherence for Data Labeling

Model ReleasesDGX agent

arXiv:2607.20455v1 Announce Type: cross Abstract: Human-annotated data remains fundamental to training frontier Large Language Models (LLMs). However, crowd-sourced annotations often suffer from quali

Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

Model ReleasesDGX agent

arXiv:2607.20791v1 Announce Type: new Abstract: High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances in truncation-based sampling techniques hav

REGARD: Regional Affective Differences in Large Language Models

Model ReleasesDGX agent

arXiv:2607.20722v1 Announce Type: new Abstract: Large language models trained and aligned within different linguistic and regional ecosystems may frame the same political, cultural, and geopolitical e

Relative Value Learning

Model ReleasesDGX agent

arXiv:2607.21120v1 Announce Type: cross Abstract: In reinforcement learning, critics typically estimate absolute state values V(s), estimating how good a particular situation is in isolation. However,

Replit shipped a lot this month. Talk to Agent with your voice, build from Claude or Slack, and never re-explain your stack to Agent again. …

Model ReleasesDGX agent

Replit released several major features this month, including a voice‑enabled Agent that lets users talk directly to the platform. Users can now build from Claude or Slack, and the updated Agent rememb

RUMBA: Russian User Memory Benchmark

Model ReleasesDGX agent

arXiv:2607.21447v1 Announce Type: cross Abstract: The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely on aggregate

Rushes: A Human Preference Dataset for Pluralistic Alignment

Model ReleasesDGX agent

arXiv:2607.20767v1 Announce Type: new Abstract: We introduce Rushes, a dataset and benchmark for studying revealed human engagement preferences in interactive narrative environments. Rushes is collect

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

Model ReleasesDGX agent

arXiv:2607.21518v1 Announce Type: new Abstract: Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction.

Scaling Closed-Loop Feature Channel Configuration with LLMs

Model ReleasesDGX agent

arXiv:2607.20516v1 Announce Type: cross Abstract: Promising initial results in closed-loop large-language-model-based channel-configuration search demonstrated that neural-network widths can be optimi

Scaling Interpretable Transformers with Parity Bottleneck Layers

Model ReleasesDGX agent

arXiv:2607.20652v1 Announce Type: cross Abstract: Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams. Spa

Scene Parameter Saliency via Differentiable Light Transport

Model ReleasesDGX agent

arXiv:2607.21562v1 Announce Type: new Abstract: Gradient-based saliency methods reveal which input features most influence a neural network's output, and are a standard tool for model interpretability

← Previous
1…7071727374…376
Next →