AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,611 results
6 Jun 2026

Evaluation of LLMs for Mathematical Formalization in Lean

Model ReleasesDGX agent

arXiv:2606.05632v1 Announce Type: new Abstract: Within the past few years, the ability of Large Language Models (LLMs) to generate formal mathematical proofs has improved drastically. We provide a com

Exploring LLMs for South Asian Music Understanding and Generation

Model ReleasesDGX agent

arXiv:2606.05522v1 Announce Type: cross Abstract: Recent advancements in Large Language Models (LLMs) have shown promising results in music understanding and generation tasks. However, existing works

Fireworks Training Platform keeps expanding. Leading US open weight model Nemotron 3 Ultra is now ready for post-training: SFT and DPO via L…

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Fireworks Training Platform keeps expanding. Leading US open weight model Nemotron 3 Ultra is now ready for post-training: SFT and DPO via LoRA or full-parameter, on the same infrastructure that serve

Four papers out recently: 1. http://political-manipulation.ai: Measures and reduces political bias in LLMs; Claude is especially biased 2. h…

Model ReleasesDGX agent

Four papers out recently: 1. http://political-manipulation.ai: Measures and reduces political bias in LLMs; Claude is especially biased 2. http://aibetrayal.com: The public can insert backdoors into A

GenTI: Benchmarking LLMs for Autonomous IDPS Rule Generation for Unseen Attacks

Model ReleasesDGX agent

arXiv:2606.05844v1 Announce Type: cross Abstract: Rule-based Intrusion Detection and Prevention Systems (IDPS) offer precise attack detection as well as mitigation, however their manually crafted, sig

Geographic Bias and Diversity in AI Evaluation

Model ReleasesDGX agent

arXiv:2606.05187v1 Announce Type: cross Abstract: Among the many challenges hindering the responsible development and deployment of AI, arguably none has faced more intense scrutiny than bias in its v

GITCO: Gated Inference-Time Context Optimization in TSFMs

Model ReleasesDGX agent

arXiv:2606.05332v1 Announce Type: new Abstract: Patch-based Time Series Foundation Models (TSFMs) suffer from context poisoning: structurally anomalous patches capture disproportionate attention and s

Goedel-Architect: Streamlining Formal Theorem Proving with Blueprint Generation and Refinement

Model ReleasesDGX agent

arXiv:2606.06468v1 Announce Type: new Abstract: We introduce Goedel-Architect, an agentic framework for formal theorem proving in Lean 4 centered on blueprint generation and refinement. A blueprint is

GuardNet: Ensemble Strategies of Shallow Neural Networks for Robust Prompt Injection and Jailbreak Detection

Model ReleasesDGX agent

arXiv:2606.05566v1 Announce Type: new Abstract: Large Language Models (LLMs) have transformed natural language processing, but they remain vulnerable to Prompt Injection (PI) and Jailbreak (JB) attack

How Far Did They Go? The Persuasive Tactics of Covert LLM Agents in a Discontinued Field Experiment

Model ReleasesDGX agent

arXiv:2606.05256v1 Announce Type: new Abstract: This study analyzes a publicly released dataset from a discontinued field experiment on Reddit's r/ChangeMyView. The intervention, conducted by unknown,

I Know What You Meme, Even If it Emerged Today: Understanding Evolving Memes through Open-World Knowledge Acquisition

Model ReleasesDGX agent

arXiv:2606.05316v1 Announce Type: new Abstract: Multimodal memes are dynamic and often require up to date background knowledge for interpretation. Existing methods often overlook such knowledge or rel

I want AI to succeed, and have been positive about many systems (Claude Code, AlphaFold, AlphaGeometry, Cicero, etc). It’s not AI that I hat…

Model ReleasesDGX agent

I want AI to succeed, and have been positive about many systems (Claude Code, AlphaFold, AlphaGeometry, Cicero, etc). It’s not AI that I hate; it’s bullshit and greed that I despise. Not my fault that

Introducing Harness-1, a 20B search agent trained with a state-externalizing harness. > frontier-level long-horizon search, rivaling Opus-4.…

Model ReleasesDGX agent

Introducing Harness-1, a 20B search agent trained with a state-externalizing harness. > frontier-level long-horizon search, rivaling Opus-4.6 and outperforming GPT-5.4 > Context-1-level cost and laten

it’s a bad day when @scaling01 goes all Gary Marcus on the media 🤣

Model ReleasesDGX agent

it’s a bad day when @scaling01 goes all Gary Marcus on the media 🤣 This is a good example of how AI news gets distorted. The example is from Tagesschau, the most watched news program in Germany. Anthr

“I've tried to find where this stops being circular and I can't.”

Model ReleasesDGX agent

“I've tried to find where this stops being circular and I can't.” 🦔Google signed a deal to pay SpaceX 920 million a month for access to 110,000 Nvidia GPUs at SpaceX data centers. The contract runs Oc

Learning What Matters: Probabilistic Task Selection via Mutual Information for Model Finetuning

Model ReleasesDGX agent

arXiv:2507.12612v3 Announce Type: replace-cross Abstract: Supervised fine-tuning performance for large language models depends strongly on how training budget is distributed across a heterogeneous set

LLMCodec: Adapting Video Codecs for Efficient Weight Compression of Large Language Models

Model ReleasesDGX agent

arXiv:2606.05861v1 Announce Type: cross Abstract: The rapid development of large language models(LLMs) has led to remarkable advances in natural language processing. However, the increasing scale of t

Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2606.06036v1 Announce Type: new Abstract: Despite recent progress, LLM agents still struggle with reasoning over long interaction histories. While current memory-augmented agents rely on a stati

Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models

Model ReleasesDGX agent

arXiv:2606.05429v1 Announce Type: new Abstract: Post-training quantization (PTQ) is critical for the efficient deployment of large language models (LLMs). Recent ultra-low-bit PTQ methods rely on rigi

Multilingual Fine-Tuning via Localized Gradient Conflict Resolution

Model ReleasesDGX agent

arXiv:2606.05613v1 Announce Type: new Abstract: The rapid evolution of Large Language Models (LLMs) has established cross-lingual versatility as a defining feature of modern systems. However, fine-tun

No Need to Train Your RDB Foundation Model

Model ReleasesDGX agent

arXiv:2602.13697v2 Announce Type: replace Abstract: Relational databases (RDBs) contain vast amounts of heterogeneous tabular information that can be exploited for predictive modeling purposes. But si

OPRD: On-Policy Representation Distillation

Model ReleasesDGX agent

arXiv:2606.06021v1 Announce Type: cross Abstract: On-policy distillation (OPD) supervises the student only in output space by matching next-token probabilities. This output-only paradigm has two limit

PC Layer: Polynomial Weight Preconditioning for Improving LLM Pre-Training

Model ReleasesDGX agent

arXiv:2606.06470v1 Announce Type: cross Abstract: We propose a preconditioning (PC) layer, a weight parameterization via polynomial preconditioner that ensures stable weight conditioning throughout LL

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows

Model ReleasesDGX agent

arXiv:2605.12376v2 Announce Type: replace Abstract: Table processing-including cleaning, transformation, augmentation, and matching-is a foundational yet error-prone stage in real-world data pipelines

PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

Model ReleasesDGX agent

arXiv:2606.05463v1 Announce Type: new Abstract: Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically perf

Retry Policy Gradients in Continuous Action Spaces

Model ReleasesDGX agent

arXiv:2606.05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that the

Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing

Model ReleasesDGX agent

arXiv:2605.04733v2 Announce Type: replace Abstract: Text-based role-playing models can imitate character styles, but often fail to capture scene atmosphere and evolving tension, which are crucial for

Reward Learning through Ranking Mean Squared Error

Model ReleasesDGX agent

arXiv:2601.09236v3 Announce Type: replace-cross Abstract: Reward design remains a significant bottleneck in applying reinforcement learning (RL) to real-world problems. A popular alternative is reward

Running Python code in a sandbox with MicroPython and WASM

Model ReleasesDGX agent

I've been experimenting with different approaches to running code in a sandbox for several years now, but my latest attempt feels like it might finally have all of the characteristics I've been lookin

Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack

Model ReleasesDGX agent

arXiv:2606.05614v1 Announce Type: new Abstract: Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and r

SagnacAssisted Enhanced OTDR for Distributed Acoustic Sensing: A Standardized Benchmark and Engineering Evaluation Framework

Model ReleasesDGX agent

arXiv:2606.05754v1 Announce Type: cross Abstract: Phase-sensitive optical time-domain reflectometry (phi-OTDR) is widely used in large-scale distributed acoustic sensing (DAS) because it provides dist

SciVisAgentSkills: Design and Evaluation of Agent Skills for Scientific Data Analysis and Visualization

Model ReleasesDGX agent

arXiv:2606.05525v1 Announce Type: new Abstract: Recent advances in agentic visualization have enabled the translation of natural language into executable scientific visualization (SciVis) workflows. W

Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation

Model ReleasesDGX agent

arXiv:2606.05241v1 Announce Type: cross Abstract: Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the

⚠️⚠️ Seismic shift ⚠️⚠️ It’s a good day to be Mistral. Nobody is going to trust an American AI company that is partly owned by the US Govern…

Model ReleasesDGX agent

⚠️⚠️ Seismic shift ⚠️⚠️ It’s a good day to be Mistral. Nobody is going to trust an American AI company that is partly owned by the US Government. Just the way the US doesn’t trust Huawei. After this m

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models

Model ReleasesDGX agent

arXiv:2606.05434v1 Announce Type: cross Abstract: Group Relative Policy Optimisation (GRPO) has emerged as an effective reinforcement-learning algorithm for aligning language models on reasoning tasks

SentinelBench: A Benchmark for Long-Running Monitoring Agents

Model ReleasesDGX agent

arXiv:2606.05342v1 Announce Type: new Abstract: AI agents are increasingly asked to carry out work that spans minutes, hours, or longer. Yet the default model of agent behavior is continuous action: i

“.. So now Google and Anthropic pay rent on the hardware Grok couldn't use, and that rent is the AI revenue story SpaceX takes public on Thu…

Model ReleasesDGX agent

“.. So now Google and Anthropic pay rent on the hardware Grok couldn't use, and that rent is the AI revenue story SpaceX takes public on Thursday.” 🦔Google signed a deal to pay SpaceX 920 million a mo

Spent more time learning deepagents from @LangChain with primary focus on integrating MCP servers with auth that doesn't conform completely …

Model ReleasesDGX agent

Spent more time learning deepagents from @LangChain with primary focus on integrating MCP servers with auth that doesn't conform completely to the OAuth standard. ( In prep of a work coming my way ) B

Starlink launches now substantially outnumber all other satellite launch sources.

Model ReleasesDGX agent

SpaceX's Starlink satellite launches have become the dominant source of orbital launches, exceeding the combined launch activity of all other satellite operators and launch providers. This milestone r

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces

Model ReleasesDGX agent

arXiv:2606.05464v1 Announce Type: new Abstract: Verifiable reward training has improved mathematical and coding reasoning, but these domains capture only part of step-by-step decision making. Many rea

Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability

Model ReleasesDGX agent

arXiv:2606.06333v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) are widely used for mechanistic interpretability in large language models, yet their formulation assigns each latent featur

Synapse: Federated Tool Routing via Typed Compendium Artifacts

Model ReleasesDGX agent

arXiv:2602.00911v2 Announce Type: replace Abstract: The unit of collaboration in federated learning determines what guarantees are even expressible. Flat units like weights, prompts, raw examples, car

Synthetic Contrastive Reasoning for Multi-Table Q&A

Model ReleasesDGX agent

arXiv:2606.05382v1 Announce Type: new Abstract: Multi-table question answering requires models to retrieve relevant evidence, link schemas, and perform compositional reasoning across relational tables

TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents

Model ReleasesDGX agent

arXiv:2606.05784v1 Announce Type: new Abstract: We identify and formally characterize credit misassignment as a systematic failure mode of GRPO in tool-augmented multimodal search agents: its uniform

The End of Software Engineering: How AI Agents Are Fundamentally Restructuring the Software Paradigm

Model ReleasesDGX agent

arXiv:2606.05608v1 Announce Type: cross Abstract: For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into s

The Gemini Pro models do not seem to be iterating anywhere near as quickly as Claude or GPT (last release was 3.1 Pro in February). Its caus…

Model ReleasesDGX agent

The Gemini Pro models do not seem to be iterating anywhere near as quickly as Claude or GPT (last release was 3.1 Pro in February). Its causing a growing performance gap between Google and the other t

TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2606.06133v1 Announce Type: cross Abstract: TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequently produ

TokenMizer: Graph-Structured Session Memory for Long-Horizon LLM Context Management

Model ReleasesDGX agent

arXiv:2606.06337v1 Announce Type: new Abstract: Large language model (LLM) deployments for long-horizon tasks face a fundamental constraint: context windows are finite while productive work sessions a

ToolChoiceConfusion: Causal Minimal Tool Filtering for Reliable LLM Agents

Model ReleasesDGX agent

arXiv:2606.06284v1 Announce Type: new Abstract: Large language model agents increasingly rely on external tools, but larger tool menus can reduce reliability and efficiency by increasing wrong-tool ca

Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation

Model ReleasesDGX agent

arXiv:2606.05403v1 Announce Type: cross Abstract: Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions. Whether they evaluate the qual

VLA-JEPA just dropped in LeRobot 🤖 What makes this model special is that it does not just learn what action to take from a given observatio…

Model ReleasesDGX agent

VLA-JEPA just dropped in LeRobot 🤖 What makes this model special is that it does not just learn what action to take from a given observation, it also leverages a JEPA world model to learn action-relev

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents

Model ReleasesDGX agent

arXiv:2606.06453v1 Announce Type: new Abstract: Sparse attention is becoming increasingly important for serving large language models (LLMs) as generation lengths continue to grow. However, deploying

What Should Agents Say? Action-state Communication for Efficient Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2606.05304v1 Announce Type: new Abstract: Multi-agent systems (MAS) built on large language models are typically organized around roles, pipelines, and turn schedules, while the content that age

When Should Memory Stay Silent: Measuring Memory-Use Boundaries in Memory-Augmented Conversational Agents

Model ReleasesDGX agent

arXiv:2606.06055v1 Announce Type: new Abstract: Long-term memory enables language model agents to support personalized interactions, but it remains unclear when available memories warrant integration

When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents

Model ReleasesDGX agent

arXiv:2606.05806v1 Announce Type: new Abstract: Existing benchmarks evaluate Tool-Integrated Reasoning (TIR) in LLMs on idealized ''happy paths'', largely overlooking real-world tool failures. We intr

Will the Agent Recuse Itself? Measuring LLM-Agent Compliance with In-Band Access-Deny Signals

Model ReleasesDGX agent

arXiv:2606.06460v1 Announce Type: cross Abstract: As autonomous LLM agents increasingly hold real credentials and operate infrastructure without a human in the loop, operators have no standard way to

Wordle 1,812 4/6 ⬛⬛⬛⬛⬛ 🟨⬛🟨⬛🟨 ⬛🟨🟨⬛🟨 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post shows a completed Wordle game (#1,812) solved in 4 attempts, displaying the colored tile feedback pattern (gray for incorrect letters, yellow for correct letters in wrong positions, green fo

WorldFly: A World-Model-Based Vision-Language-Action Model for UAV Navigation

Model ReleasesDGX agent

arXiv:2606.06147v1 Announce Type: new Abstract: End-to-end Vision-Language-Action (VLA) models have shown promise in UAV navigation. However, existing approaches typically rely on historical observati

You can’t let this happen, @DavidSacks, cc @elonmusk. It’s WWWIII but with AI. Nobody wins.

Model ReleasesDGX agent

You can’t let this happen, @DavidSacks, cc @elonmusk. It’s WWWIII but with AI. Nobody wins. ⚠️⚠️ Seismic shift ⚠️⚠️ It’s a good day to be Mistral. Nobody is going to trust an American AI company that

Your margin is my opportunity: AI version… The biggest surprise of 2026 is that the capability gap between the best open-weight/source model…

Model ReleasesDGX agent

Your margin is my opportunity: AI version… The biggest surprise of 2026 is that the capability gap between the best open-weight/source models and the best closed models has narrowed much faster than t

← Previous
1…165166167168169…377
Next →