AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,606
  • Agents7,269
  • Applications5,200
  • Concepts5
  • Hardware1,756
  • Industry6,099
  • Local Ai4,731
  • Model Releases22,585
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,606Total entries
1Added by human
84,605Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,585 results
Model Releases

An Empirical Study of Security Calibration in Large Language Models for Code

DGX agent

arXiv:2606.31159v1 Announce Type: cross Abstract: Large Language Models (LLMs) are rapidly transforming software development, yet their use in security-critical contexts raises a key question: do mode

model-releasesarxiv-cs-lg
1 Jul 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

An Executable Benchmarking Suite for Tool-Using Agents

DGX agent

arXiv:2605.11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often con

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Anthropic launches Claude Sonnet 5 AI model with coding, safety upgrades as Fable and Mythos controls lifted

DGX agent

Anthropic PBC today debuted Claude Sonnet 5, a midrange large language model that outperforms its predecessor in several areas. The LLM will be the default option in the consumer tiers of the company’

model-releasessiliconangle
1 Jul 2026
Model Releases

Anthropic says Fable 5 will be available via usage credits for Claude users from July 7, and is working with partners to draft an AI jailbreak severity standard (Anthropic)

DGX agent

Anthropic: Anthropic says Fable 5 will be available via usage credits for Claude users from July 7, and is working with partners to draft an AI jailbreak severity standard — On Friday, June 12, the US

model-releasestechmeme
1 Jul 2026
Model Releases

Anthropic says it is rolling back a covert Claude Code tracking feature to identify users based in China or affiliated with Chinese AI labs, after backlash (Juro Osawa/The Information)

DGX agent

Juro Osawa / The Information: Anthropic says it is rolling back a covert Claude Code tracking feature to identify users based in China or affiliated with Chinese AI labs, after backlash — Anthropic is

model-releasestechmeme
1 Jul 2026
Model Releases

Anthropic says 'some routine tasks like coding and debugging' on Fable 5 'will fall back to Opus 4.8' in 'the near term' as it works to 'reduce false positives' (@anthropicai)

DGX agent

@anthropicai: Anthropic says “some routine tasks like coding and debugging” on Fable 5 “will fall back to Opus 4.8” in “the near term” as it works to “reduce false positives” — Claude Fable 5 will be

model-releasestechmeme
1 Jul 2026
Model Releases

ARC-AGI-3 is built different, it has dumbfounded almost all regular attempts so far because it's so much harder than anything that came befo…

DGX agent

ARC-AGI-3 is built different, it has dumbfounded almost all regular attempts so far because it's so much harder than anything that came before. It has no rules, it's agentic and has no explicit goals,

model-releasesfrancois-chollet--x
1 Jul 2026
Model Releases

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

DGX agent

arXiv:2606.31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) model

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Armadin details full sandbox escape in Claude Cowork but Anthropic disputes risk

DGX agent

Security researchers at Armadin Inc. today detailed an attack chain that runs arbitrary commands as root inside the sandbox behind Anthropic PBC’s Claude Cowork, escaping the isolation layer, with a s

model-releasessiliconangle
1 Jul 2026
Model Releases

Artificial Intelligence in Sports: Insights from a Quantitative Survey among Sports Students in Germany about their Perceptions, Expectations, and Concerns regarding the Use of AI Tools

DGX agent

arXiv:2503.05785v2 Announce Type: replace-cross Abstract: Generative Artificial Intelligence (AI) tools such as ChatGPT, Copilot, or Gemini have a crucial impact on academic research and teaching. Emp

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

As generative AI tools continue to evolve, we believe it's more important than ever to know what's AI-generated and what isn't. That’s why @…

DGX agent

As generative AI tools continue to evolve, we believe it's more important than ever to know what's AI-generated and what isn't. That’s why @GoogleDeepMind launched SynthID in 2023—a technology that ad

model-releasesgoogle-ai--x
1 Jul 2026
Model Releases

AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

DGX agent

arXiv:2606.31551v1 Announce Type: new Abstract: Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents become increasingly capable at software

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

AxDafny: Agentic Verified Code Generation in Dafny

DGX agent

arXiv:2606.32007v1 Announce Type: new Abstract: We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

DGX agent

arXiv:2606.30850v1 Announce Type: new Abstract: Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic unce

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Benchmarking Large Language Models on Floating-Point Error Classification

DGX agent

arXiv:2606.31308v1 Announce Type: new Abstract: This paper investigates the capability of Large Language Models (LLMs) to detect and classify floating-point errors statically in software code. We intr

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Best Claude use case ever: learning to use Microsoft Teams for first time 🙃🤣 - from @_catwu at AIE w/ @swyx & @trq212

DGX agent

This post humorously describes using Claude AI as a learning tool to understand Microsoft Teams functionality for first-time users, shared during an AI Engineer event discussion. The post suggests Cla

model-releasesswyx--x
1 Jul 2026
Model Releases

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

DGX agent

arXiv:2606.31338v1 Announce Type: cross Abstract: Recent music audio-language models achieve high accuracy on instrument question-answering benchmarks, but it remains unclear whether this reflects rob

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Beyond Clean Text: Evaluating Encoder and Decoder Robustness for Bangla Event Detection in Noisy Text

DGX agent

arXiv:2606.30914v1 Announce Type: new Abstract: Event detection (ED) systems are typically evaluated on clean, curated text, leaving their robustness to real-world noise largely unexplored, particular

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization

DGX agent

arXiv:2606.31002v1 Announce Type: new Abstract: Theorem-proving benchmarks evaluate proof search against fixed formal statements, but natural-language-to-Lean formalization must generate the formal st

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Beyond expert users: agents should help users construct preferences, not just elicit them

DGX agent

arXiv:2606.30863v1 Announce Type: new Abstract: Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

DGX agent

arXiv:2606.31169v1 Announce Type: new Abstract: Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

Beyond Static Prompts: Building Scale-Proof, Polymorphic Multi-Agent Systems with Google's ADK

DGX agent

As enterprise generative AI transitions from simple, conversational chatbots to autonomous multi-agent workflows, developers face a critical bottleneck: scale. In a production environment, an enterpri

model-releasesgoogle-cloud-ai
1 Jul 2026
Model Releases

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

DGX agent

arXiv:2606.22723v2 Announce Type: replace Abstract: Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended,

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Bridging Scientific Heritage: An Arabic--Russian Parallel Corpus and LLM Benchmark for Sustainable Knowledge Transfer

DGX agent

arXiv:2606.30943v1 Announce Type: new Abstract: Russian and Arabic are among the major languages of scientific communication. Language barriers impede the exchange of research results between these co

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Building an ASR Solution for Training and Assessing Children's Reading

DGX agent

arXiv:2606.31508v1 Announce Type: new Abstract: Automatic speech recognition for children's reading remains underdeveloped for most African languages, including Bambara, despite its potential value fo

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?

DGX agent

arXiv:2606.31371v1 Announce Type: cross Abstract: When large language model (LLM) agents adapt their behavior through evaluator feedback, systematic evaluator biases propagate into the agent's learned

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models

DGX agent

arXiv:2606.31630v1 Announce Type: new Abstract: Language models increasingly write probabilistic programs (in NumPyro, Stan, or Pyro), but a program that compiles, runs, and passes every unit test can

model-releasesarxiv-cs-lg
1 Jul 2026
Model Releases

Can Physician Expertise Improve Machine Learning Identification of Delirium?

DGX agent

arXiv:2606.30651v1 Announce Type: cross Abstract: Delirium is common in hospitalized patients and is often missed in routine care. We present a user-centered interactive machine learning (UC-iML) fram

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

@_catwu @simonw and I will be doing a fireside chat about 'This year in Claude' from 12:30pm-1:30pm at AIE in Expo Stage 2. We'll be coverin…

DGX agent

@_catwu @simonw and I will be doing a fireside chat about 'This year in Claude' from 12:30pm-1:30pm at AIE in Expo Stage 2. We'll be covering a really wide range of topics and I think it will be reall

model-releasessimon-willison--x
1 Jul 2026
Model Releases

CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes

DGX agent

arXiv:2606.31435v1 Announce Type: new Abstract: Data refinement involves executing multi-step recipes over evolving text states, where both composition and execution order of processing operators dete

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield

DGX agent

arXiv:2606.31796v1 Announce Type: cross Abstract: We study three complementary techniques for training compute-efficient language models. (1) Selective supervision and per-token efficiency. Selective

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Chinese robot maker UBTech launches U1, a line of humanoid robots for personal companionship with lifelike silicone skin and emotional AI, priced from $17,650 (Minxiao Chang/South China Morning Post)

DGX agent

Minxiao Chang / South China Morning Post: Chinese robot maker UBTech launches U1, a line of humanoid robots for personal companionship with lifelike silicone skin and emotional AI, priced from $17,650

model-releasestechmeme
1 Jul 2026
Model Releases

Citation Discipline in Spec-Driven Development: A Cross-Model Empirical Study of Output Determinism and Automated Hallucination Detection in LLM-Generated Code

DGX agent

arXiv:2606.30689v1 Announce Type: cross Abstract: Spec-Driven Development (SDD) frameworks guide Large Language Model (LLM)-powered code generation through formal specifications, yet they differ funda

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Claude Fable 5 access restored on AI Gateway

DGX agent

Vercel has restored access to Claude Fable 5 on its AI Gateway service, allowing developers to integrate the model into their applications through Vercel's platform. This restoration enables continued

model-releasesvercel-blog
1 Jul 2026
Model Releases

Claude Fable 5 is available again in Cursor. It leads all models on CursorBench, but is the most expensive per task.

DGX agent

Claude Fable 5 has been re-enabled as an available model option within Cursor, where it achieves the highest performance scores on CursorBench among all supported models. However, it comes with the hi

model-releasescursor--x
1 Jul 2026
Model Releases

Claude Fable 5 is available in Devin. You can use Fable 5 in Devin Cloud’s Ultra agent, our smartest and most capable agent, which excels at…

DGX agent

Claude Fable 5 is available in Devin. You can use Fable 5 in Devin Cloud’s Ultra agent, our smartest and most capable agent, which excels at long-horizon tasks and debugging. Claude Fable 5 is also av

model-releasescognition-ai--x
1 Jul 2026
Model Releases

Claude Fable 5 is once again available in Computer as an orchestrator model.

DGX agent

Claude Fable 5 has been re-enabled as an orchestrator model option within Computer, Anthropic's AI system for automated task execution. This announcement from Perplexity indicates that users can again

model-releasesperplexity--x
1 Jul 2026
Model Releases

Claude Fable 5 will be available again globally tomorrow. After a series of productive conversations with the US government, we're redeployi…

DGX agent

Claude Fable 5 will be available again globally tomorrow. After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to target and blo

model-releasesjerry-liu--x
1 Jul 2026
Model Releases

ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents

DGX agent

arXiv:2606.31174v1 Announce Type: new Abstract: Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized sub

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Clinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question Answering

DGX agent

arXiv:2606.31432v1 Announce Type: new Abstract: Medical multiple-choice question answering requires parameter-efficient adaptation across heterogeneous knowledge domains and reasoning operations. A me

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Coarsening Bias from Variable Discretization in Causal Functionals

DGX agent

arXiv:2602.22083v2 Announce Type: replace-cross Abstract: Causal identification functionals often require integration over conditional densities of continuous variables, such as those arising in nonpa

model-releasesarxiv-cs-lg
1 Jul 2026
Model Releases

CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts

DGX agent

arXiv:2606.31986v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning has enabled multi-modal large language models (MLLMs) to tackle complex visual reasoning tasks by generating explicit i

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm

DGX agent

arXiv:2606.13239v2 Announce Type: replace-cross Abstract: Existing computer-use agents remain fundamentally limited in professional software manipulation: GUI-based agents suffer from fragile visual g

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Contextual Slate GLM Bandits with Limited Adaptivity

DGX agent

arXiv:2606.31449v1 Announce Type: new Abstract: We investigate the contextual slate bandit problem with generalized linear rewards under limited adaptivity. At each round, the learner is presented wit

model-releasesarxiv-cs-lg
1 Jul 2026
Model Releases

CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization

DGX agent

arXiv:2606.31219v1 Announce Type: new Abstract: Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However,

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

Cross-agent feedback loops are incredibly effective -- for a reason. Check out what @leon2mcp and team at @Bloome_im are building in this sp…

DGX agent

Cross-agent feedback loops are incredibly effective -- for a reason. Check out what @leon2mcp and team at @Bloome_im are building in this space: http://bloome.im Bloome lets you pull Claude, ChatGPT,

model-releasesfrancois-chollet--x
1 Jul 2026
Model Releases

Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian

DGX agent

arXiv:2606.31718v1 Announce Type: cross Abstract: Relation extraction (RE) for low-resource languages is typically constrained by the lack of annotated corpora. We investigate the feasibility of cross

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market

DGX agent

arXiv:2606.31461v1 Announce Type: new Abstract: Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and platform rules. T

model-releasesarxiv-cs-ai
1 Jul 2026
← Previous
1…140141142143144…471
Next →