AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “anthropic”

GridTimelineEvolution
49+ results
Model Releases

How Well Do Models Follow Their Constitutions?

DGX agent

arXiv:2605.24229v1 Announce Type: new Abstract: Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's

model-releasesarxiv-cs-ai
26 May 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Authority Signals in Claude AI Health Citations: A Descriptive Analysis Using the Authority Signals Framework

DGX agent

arXiv:2605.23921v1 Announce Type: cross Abstract: This study seeks to determine the authority signals used by Anthropic's Claude AI in its presentation of sources when answering consumer health questi

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Benchmarking Mythos-Linked Bug Rediscovery

DGX agent

arXiv:2605.17416v1 Announce Type: cross Abstract: Anthropic's April 2026 Mythos materials combine benchmark claims with concrete bug-finding stories across OpenBSD, FreeBSD, Linux, FFmpeg, and browser

model-releasesarxiv-cs-ai
19 May 2026
Model Releases

Authoring Agent Skills: A Software-Engineering Approach

DGX agent

arXiv:2607.25032v1 Announce Type: cross Abstract: Agent Skills are an emerging way to extend large language model agents with reusable procedural knowledge that the agent loads on demand. Anthropic in

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

DGX agent

arXiv:2607.06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI

DGX agent

arXiv:2607.01418v1 Announce Type: cross Abstract: Organizations rolling out agentic command line tools like Anthropic's Claude Code and GitHub's Copilot CLI need to know who will try them, who will ke

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

MCP Server Architecture Patterns for LLM-Integrated Applications

DGX agent

arXiv:2606.30317v1 Announce Type: cross Abstract: The Model Context Protocol (MCP), introduced by Anthropic in November 2024, defines a standardized interface for connecting large language models (LLM

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO

DGX agent

arXiv:2606.23032v2 Announce Type: replace Abstract: Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language model

model-releasesarxiv-cs-ai
24 Jun 2026
Model Releases

Chatbots Output Meaningful (but Problematic) Language

DGX agent

arXiv:2606.02973v1 Announce Type: new Abstract: Are utterances by AI chatbots meaningful? Concretely, if a user asks, say, Anthropic's agent Claude, 'What is the capital of Spain?' and Claude answers,

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope

DGX agent

arXiv:2605.28916v1 Announce Type: cross Abstract: We report a comparison of two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a si

model-releasesarxiv-cs-ai
29 May 2026
Applications

Persona Conditioning of Brand Recommendations in Retrieval-Augmented Commercial Chat: A Prominence-Stratified Cross-Provider Audit

DGX agent

arXiv:2605.30207v1 Announce Type: new Abstract: The same prompt -- 'best CRM software' -- reaches AI assistants from buyers in widely different contexts: a solo founder, an enterprise VP, a UK SMB own

applicationsarxiv-cs-ai
29 May 2026
Agents

A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology

DGX agent

arXiv:2605.13850v1 Announce Type: new Abstract: Existing frameworks for LLM-based agent architectures describe systems from a single perspective: industry guides (Anthropic, Google, LangChain) focus o

agentsarxiv-cs-ai
15 May 2026
Model Releases

TSCG: Deterministic Tool-Schema Compilation for Agentic LLM Deployments

DGX agent

arXiv:2605.04107v1 Announce Type: cross Abstract: Production agent frameworks (OpenAI Function Calling, Anthropic Tool Use, MCP) transmit tool schemas as JSON, a format designed for machine parsing, n

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification

DGX agent

arXiv:2604.01687v2 Announce Type: replace Abstract: Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address. A tool

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?

DGX agent

arXiv:2503.18018v2 Announce Type: replace Abstract: We demonstrate that large language models' (LLMs) mathematical reasoning is culturally sensitive: testing 14 models from Anthropic, OpenAI, Google,

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

DGX agent

arXiv:2607.23386v1 Announce Type: new Abstract: We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images

DGX agent

arXiv:2607.21672v1 Announce Type: cross Abstract: Long source-code contexts consume many text tokens, motivating the proposal to render code as images for vision-language models. Recent work asks whet

model-releasesarxiv-cs-cv
27 Jul 2026
Model Releases

ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems

DGX agent

arXiv:2607.19430v1 Announce Type: cross Abstract: Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel thr

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

Scaffold Effects on GAIA: A Controlled Comparison

DGX agent

arXiv:2606.08529v1 Announce Type: new Abstract: Published agent capability scores conflate what a model can do with what its scaffold lets it do, and the magnitude of this elicitation gap is not well

model-releasesarxiv-cs-ai
9 Jun 2026
Model Releases

ASE-26: a curriculum for agentic software engineering as a discipline

DGX agent

arXiv:2606.01152v1 Announce Type: cross Abstract: The work of a professional software engineer has begun to consist, increasingly, of directing agents rather than writing code, and the empirical evide

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AMEL: Accumulated Message Effects on LLM Judgments

DGX agent

arXiv:2605.22714v1 Announce Type: cross Abstract: Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing th

model-releasesarxiv-cs-cl
22 May 2026
Model Releases

Position: AI Security Policy Should Target Systems, Not Models

DGX agent

arXiv:2605.09504v1 Announce Type: cross Abstract: We present swarm-attack, an open-source adversarial testing framework in which multiple lightweight LLM agents coordinate through shared memory, paral

model-releasesarxiv-cs-ai
12 May 2026
Agents

GitSkills: A Dataset of Agent Skills on GitHub

DGX agent

arXiv:2608.10906v1 Announce Type: cross Abstract: An agent skill is a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and reference fi

agentsarxiv-cs-ai
12 Aug 2026
Model Releases

Automating Deception: Scalable Multi-Turn LLM Jailbreaks

DGX agent

arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Can Open-Weight Models Compete on Financial Text Comprehension?

DGX agent

arXiv:2608.08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability

model-releasesarxiv-cs-ai
11 Aug 2026
Safety

Stealing Reasoning Traces from Proprietary LLM APIs

DGX agent

arXiv:2608.09867v1 Announce Type: cross Abstract: Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and lim

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM Evaluation

DGX agent

arXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model 'value dispositions,' and a concentration/ex

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation

DGX agent

arXiv:2608.06955v1 Announce Type: new Abstract: Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs sys

model-releasesarxiv-cs-ai
10 Aug 2026
Research

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

DGX agent

arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m

researcharxiv-cs-cl
5 Aug 2026
Model Releases

How Closely Do LLM Reviews Align with Human Peer Review?

DGX agent

arXiv:2608.03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation

DGX agent

arXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or

model-releasesarxiv-cs-cl
31 Jul 2026
Safety

Constitutional Midtraining: Content Presence Drives Alignment Gains

DGX agent

arXiv:2607.26654v1 Announce Type: new Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce

safetyarxiv-cs-cl
30 Jul 2026
Model Releases

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

DGX agent

arXiv:2607.26062v1 Announce Type: cross Abstract: Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intell

model-releasesarxiv-cs-cl
30 Jul 2026
Safety

OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

DGX agent

arXiv:2607.26981v1 Announce Type: new Abstract: Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a syste

safetyarxiv-cs-cl
30 Jul 2026
Model Releases

Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation

DGX agent

arXiv:2607.22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instruction-tuni

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model

DGX agent

arXiv:2607.22951v1 Announce Type: cross Abstract: Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified ope

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Case study: solving P-99 with LPTP and an LLM

DGX agent

arXiv:2607.21196v1 Announce Type: cross Abstract: Ninety-Nine Prolog Problems (P-99) is a famous set of Prolog exercises. We solved the first thirty three just by prompting an LLM (Large Language Mode

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

DGX agent

arXiv:2607.20759v1 Announce Type: cross Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with au

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation

DGX agent

arXiv:2607.15434v3 Announce Type: replace-cross Abstract: Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets

DGX agent

arXiv:2607.19967v1 Announce Type: cross Abstract: Shippers are beginning to delegate carrier selection to large language model (LLM) agents. We ask what such delegation does to a freight matching mark

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

Context Graphs for Proactive Enterprise Agents

DGX agent

arXiv:2607.07721v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) and agentic frameworks have advanced enterprise AI considerably, yet agents remain fundamentally reactive: they wai

model-releasesarxiv-cs-ai
10 Jul 2026
Local Ai

SPL: Orchestrating Workflows with Declarative Deterministic-Probabilistic Composition

DGX agent

arXiv:2607.07727v1 Announce Type: cross Abstract: We present SPL (Structured Prompt Language), a declarative language that composes deterministic and probabilistic computation modes in a single specif

local-aiarxiv-cs-cl
10 Jul 2026
Safety

The Jagged Global Economy: Frontier AI Unevenly Exposes National Economies

DGX agent

arXiv:2607.05404v1 Announce Type: cross Abstract: Frontier AI's labor-market effects matter to workers, firms, and policymakers, but current evidence generally comes from a handful of high-income econ

safetyarxiv-cs-ai
8 Jul 2026
Model Releases

meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysis

DGX agent

arXiv:2606.28363v1 Announce Type: cross Abstract: Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that integrates th

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates

DGX agent

arXiv:2606.30085v1 Announce Type: new Abstract: Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produc

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation

DGX agent

arXiv:2606.26116v1 Announce Type: cross Abstract: A brand whose customers use both ChatGPT and Claude for product recommendations faces a strategic choice: a single optimization playbook, or one per p

model-releasesarxiv-cs-ai
26 Jun 2026
Model Releases

Generative AI and Copyright Infringement: A Legal-Technical Analysis of AI Music Generation Systems Under 17 U.S.C. Title 17

DGX agent

arXiv:2606.26111v1 Announce Type: cross Abstract: Generative artificial intelligence (GenAI) has enabled users to synthesize music with text prompts, combining copyrighted lyrics, AI-composed melodies

model-releasesarxiv-cs-ai
26 Jun 2026
Agents

AgentRivet: an automated system for producing Rivet routines from journal publications

DGX agent

arXiv:2606.13535v3 Announce Type: replace-cross Abstract: Particle physics collider experiments provide Rivet routines as part of the analysis preservation strategy for model-independent measurements.

agentsarxiv-cs-ai
24 Jun 2026
← Previous
1
Next →
114 results
← Previous
123
Next →