AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
Human
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “anthropic”

GridTimelineEvolution
61+ results
26 May 2026

How Well Do Models Follow Their Constitutions?

Model ReleasesDGX agent

arXiv:2605.24229v1 Announce Type: new Abstract: Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's

Authority Signals in Claude AI Health Citations: A Descriptive Analysis Using the Authority Signals Framework

Model ReleasesDGX agent

arXiv:2605.23921v1 Announce Type: cross Abstract: This study seeks to determine the authority signals used by Anthropic's Claude AI in its presentation of sources when answering consumer health questi

19 May 2026

Benchmarking Mythos-Linked Bug Rediscovery

DGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2605.17416v1 Announce Type: cross Abstract: Anthropic's April 2026 Mythos materials combine benchmark claims with concrete bug-finding stories across OpenBSD, FreeBSD, Linux, FFmpeg, and browser

29 Jul 2026

Authoring Agent Skills: A Software-Engineering Approach

Model ReleasesDGX agent

arXiv:2607.25032v1 Announce Type: cross Abstract: Agent Skills are an emerging way to extend large language model agents with reusable procedural knowledge that the agent loads on demand. Anthropic in

9 Jul 2026

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

Model ReleasesDGX agent

arXiv:2607.06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic

3 Jul 2026

Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

Model ReleasesDGX agent

arXiv:2607.01418v1 Announce Type: cross Abstract: Organizations rolling out agentic command line tools like Anthropic's Claude Code and GitHub's Copilot CLI need to know who will try them, who will ke

30 Jun 2026

MCP Server Architecture Patterns for LLM-Integrated Applications

Model ReleasesDGX agent

arXiv:2606.30317v1 Announce Type: cross Abstract: The Model Context Protocol (MCP), introduced by Anthropic in November 2024, defines a standardized interface for connecting large language models (LLM

meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysis

Model ReleasesDGX agent

arXiv:2606.28363v1 Announce Type: cross Abstract: Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that integrates th

Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates

Model ReleasesDGX agent

arXiv:2606.30085v1 Announce Type: new Abstract: Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produc

24 Jun 2026

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO

Model ReleasesDGX agent

arXiv:2606.23032v2 Announce Type: replace Abstract: Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language model

AgentRivet: an automated system for producing Rivet routines from journal publications

AgentsDGX agent

arXiv:2606.13535v3 Announce Type: replace-cross Abstract: Particle physics collider experiments provide Rivet routines as part of the analysis preservation strategy for model-independent measurements.

Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical Assessment

Model ReleasesDGX agent

arXiv:2606.24331v1 Announce Type: new Abstract: Transformer-based language models have become the default substrate for natural language processing and the pace of new releases has made it hard for pr

3 Jun 2026

Chatbots Output Meaningful (but Problematic) Language

Model ReleasesDGX agent

arXiv:2606.02973v1 Announce Type: new Abstract: Are utterances by AI chatbots meaningful? Concretely, if a user asks, say, Anthropic's agent Claude, 'What is the capital of Spain?' and Claude answers,

LAP: An Agent-to-Instrument Protocol for Autonomous Science

SafetyDGX agent

arXiv:2606.03755v1 Announce Type: new Abstract: Autonomous science is moving from demonstration to infrastructure. Large language model agents now plan experiments, and self-driving laboratories execu

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

Model ReleasesDGX agent

arXiv:2603.04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- se

29 May 2026

First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope

Model ReleasesDGX agent

arXiv:2605.28916v1 Announce Type: cross Abstract: We report a comparison of two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a si

Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit

ApplicationsDGX agent

arXiv:2605.30207v1 Announce Type: new Abstract: The same prompt -- 'best CRM software' -- reaches AI assistants from buyers in widely different contexts: a solo founder, an enterprise VP, a UK SMB own

Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes

Model ReleasesDGX agent

arXiv:2605.28965v1 Announce Type: new Abstract: Linking free-text phenotype descriptions to ontology terms, typically referred to as phenotype annotation, is essential for the cross-study integration

How Reliable Are AI Attackers Against a Fixed Vulnerable Target? A 400-Run Empirical Study of LLM Penetration Testing Consistency

Model ReleasesDGX agent

arXiv:2605.30096v1 Announce Type: cross Abstract: Large language models (LLMs) can autonomously conduct multi-stage cyber attacks, but the consistency of their offensive behavior under repeated trials

15 May 2026

A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology

AgentsDGX agent

arXiv:2605.13850v1 Announce Type: new Abstract: Existing frameworks for LLM-based agent architectures describe systems from a single perspective: industry guides (Anthropic, Google, LangChain) focus o

7 May 2026

TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments

Model ReleasesDGX agent

arXiv:2605.04107v1 Announce Type: cross Abstract: Production agent frameworks (OpenAI Function Calling, Anthropic Tool Use, MCP) transmit tool schemas as JSON, a format designed for machine parsing, n

14 Apr 2026

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

Model ReleasesDGX agent

arXiv:2604.01687v2 Announce Type: replace Abstract: Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool

10 Apr 2026

Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?

Model ReleasesDGX agent

arXiv:2503.18018v2 Announce Type: replace Abstract: We demonstrate that large language models' (LLMs) mathematical reasoning is culturally sensitive: testing 14 models from Anthropic, OpenAI, Google,

28 Jul 2026

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

Model ReleasesDGX agent

arXiv:2607.23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

Model ReleasesDGX agent

arXiv:2607.22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuni

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

Model ReleasesDGX agent

arXiv:2607.22951v1 Announce Type: cross Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified ope

27 Jul 2026

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images

Model ReleasesDGX agent

arXiv:2607.21672v1 Announce Type: cross Abstract: Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whet

23 Jul 2026

ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2607.19430v1 Announce Type: cross Abstract: Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel thr

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

Model ReleasesDGX agent

arXiv:2607.15434v3 Announce Type: replace-cross Abstract: Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome

When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets

Model ReleasesDGX agent

arXiv:2607.19967v1 Announce Type: cross Abstract: Shippers are beginning to delegate carrier selection to large language model (LLM) agents. We ask what such delegation does to a freight matching mark

9 Jun 2026

Scaffold Effects on GAIA: A Controlled Comparison

Model ReleasesDGX agent

arXiv:2606.08529v1 Announce Type: new Abstract: Published agent capability scores conflate what a model can do with what its scaffold lets it do, and the magnitude of this elicitation gap is not well

Adversarial Robustness of Activation Steering in Large Language Models

ApplicationsDGX agent

arXiv:2606.07696v1 Announce Type: cross Abstract: Activation steering has become a popular training-free method to control LLM behavior by injecting precomputed direction vectors into the model's resi

2 Jun 2026

ASE-26: a curriculum for agentic software engineering as a discipline

Model ReleasesDGX agent

arXiv:2606.01152v1 Announce Type: cross Abstract: The work of a professional software engineer has begun to consist, increasingly, of directing agents rather than writing code, and the empirical evide

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

Model ReleasesDGX agent

arXiv:2606.02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, S

ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL

SafetyDGX agent

arXiv:2606.01619v1 Announce Type: new Abstract: Agentic reinforcement learning (RL) enables LLM agents to improve continuously from environment rewards, yet the resulting policies do not systematicall

22 May 2026

AMEL: Accumulated Message Effects on LLM Judgments

Model ReleasesDGX agent

arXiv:2605.22714v1 Announce Type: cross Abstract: Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing th

12 May 2026

Position: AI Security Policy Should Target Systems, Not Models

Model ReleasesDGX agent

arXiv:2605.09504v1 Announce Type: cross Abstract: We present swarm-attack, an open-source adversarial testing framework in which multiple lightweight LLM agents coordinate through shared memory, paral

12 Aug 2026

GitSkills: A Dataset of Agent Skills on GitHub

AgentsDGX agent

arXiv:2608.10906v1 Announce Type: cross Abstract: An agent skill is a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and reference fi

11 Aug 2026

Automating Deception: Scalable Multi-Turn LLM Jailbreaks

Model ReleasesDGX agent

arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t

Can Open-Weight Models Compete on Financial Text Comprehension?

Model ReleasesDGX agent

arXiv:2608.08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability

Stealing Reasoning Traces from Proprietary LLM APIs

SafetyDGX agent

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation

Model ReleasesDGX agent

arXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model 'value dispositions,' and a concentration/ex

10 Aug 2026

Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

Model ReleasesDGX agent

arXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys

5 Aug 2026

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

ResearchDGX agent

arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m

How Closely Do LLM Reviews Align with Human Peer Review?

Model ReleasesDGX agent

arXiv:2608.03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers

31 Jul 2026

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or

30 Jul 2026

Constitutional Midtraining: Content Presence Drives Alignment Gains

SafetyDGX agent

arXiv:2607.26654v1 Announce Type: new Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

Model ReleasesDGX agent

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

SafetyDGX agent

arXiv:2607.26981v1 Announce Type: new Abstract: Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a syste

24 Jul 2026

Case study: solving P-99 with LPTP and an LLM

Model ReleasesDGX agent

arXiv:2607.21196v1 Announce Type: cross Abstract: Ninety-Nine Prolog Problems (P-99) is a famous set of Prolog exercises. We solved the first thirty three just by prompting an LLM (Large Language Mode

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

Model ReleasesDGX agent

arXiv:2607.20759v1 Announce Type: cross Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with au

10 Jul 2026

Context Graphs for Proactive Enterprise Agents

Model ReleasesDGX agent

arXiv:2607.07721v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) and agentic frameworks have advanced enterprise AI considerably, yet agents remain fundamentally reactive: they wai

SPL: Orchestrating Workflows with Declarative Deterministic-Probabilistic Composition

Local AiDGX agent

arXiv:2607.07727v1 Announce Type: cross Abstract: We present SPL (Structured Prompt Language), a declarative language that composes deterministic and probabilistic computation modes in a single specif

8 Jul 2026

The Jagged Global Economy: Frontier AI Unevenly Exposes National Economies

SafetyDGX agent

arXiv:2607.05404v1 Announce Type: cross Abstract: Frontier AI's labor-market effects matter to workers, firms, and policymakers, but current evidence generally comes from a handful of high-income econ

26 Jun 2026

Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation

Model ReleasesDGX agent

arXiv:2606.26116v1 Announce Type: cross Abstract: A brand whose customers use both ChatGPT and Claude for product recommendations faces a strategic choice: a single optimization playbook, or one per p

Generative AI and Copyright Infringement: A Legal-Technical Analysis of AI Music Generation Systems Under 17 U.S.C. Title 17

Model ReleasesDGX agent

arXiv:2606.26111v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has enabled users to synthesize music with text prompts, combining copyrighted lyrics, AI-composed melodies

10 Jun 2026

Constructing coherent spatial memory in LLM agents through graph rectification

Model ReleasesDGX agent

arXiv:2510.04195v2 Announce Type: replace Abstract: Given a map description through global traversal navigation instructions, an LLM can often infer the implicit spatial layout and answer user queries

8 Jun 2026

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling

SafetyDGX agent

arXiv:2507.06419v3 Announce Type: replace Abstract: Reward modeling (RM), which captures human preferences to align large language models (LLMs), is increasingly employed in tasks such as model finetu

4 Jun 2026

Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery

AgentsDGX agent

arXiv:2606.05037v1 Announce Type: cross Abstract: When an AI agent calls an API and hits a validation error, it needs more than what went wrong -- it needs what to do next. A self-reflective API retur

28 May 2026

Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline

Model ReleasesDGX agent

arXiv:2605.27440v1 Announce Type: cross Abstract: Small changes to how a buyer phrases a question -- 'best CRM' vs 'top CRM' vs 'best CRM for a SaaS startup' -- produce substantially different brand r

← Previous
1
Next →
114 results
← Previous
12
Next →