AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,084 results
Model Releases

Anthropic says new Claude models in the EU will add watermarks to text and C2PA metadata to files, to comply with the EU AI Act, and it will update older models (Thomas Claburn/The Register)

DGX agent

Thomas Claburn / The Register: Anthropic says new Claude models in the EU will add watermarks to text and C2PA metadata to files, to comply with the EU AI Act, and it will update older models — EU rul

model-releasestechmeme
11 Aug 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Anthropic to start watermarking Claude-generated text, images

DGX agent

Anthropic PBC has announced plans to embed an invisible watermark in text and images generated by Claude. The Register reported the change today, citing a help desk article published on Monday. It app

model-releasessiliconangle
11 Aug 2026
Model Releases

AnyCamVLA: Zero-Shot Camera Adaptation for Viewpoint Robust Vision-Language-Action Models

DGX agent

arXiv:2603.05868v2 Announce Type: replace Abstract: Despite remarkable progress in Vision-Language-Action models (VLAs) for robot manipulation, these large pre-trained models require fine-tuning to be

model-releasesarxiv-cs-ro
11 Aug 2026
Model Releases

Anyone Using (Koreas) 'Solar Open 2' (250B, 15B) Model?

DGX agent

I just heard of this model. Seems to be a competitor to DeepSeek V4 Flash. About the same size and active parameters. Anyone tested it compared to V4 Flash? Link: https://huggingface.co/upstage/Solar-

model-releasesr-localllama
11 Aug 2026
Model Releases

APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain

DGX agent

arXiv:2608.08059v1 Announce Type: new Abstract: Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, espe

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

AQUA20: A Benchmark Dataset for Underwater Species Classification under Challenging Conditions

DGX agent

arXiv:2506.17455v3 Announce Type: replace Abstract: Robust visual recognition in underwater environments remains a significant challenge due to complex distortions such as turbidity, low illumination,

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

ARC: Augmented-Rank Conformalization for Changepoint Localization --- Finite-Sample Validity and Distribution-Robust Efficiency

DGX agent

arXiv:2608.08424v1 Announce Type: cross Abstract: Conformal changepoint localization turns any score into a confidence set for the changepoint with finite-sample coverage. Coverage is universal; effic

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

ASPaeroFlow: Decomposition Heuristics for Joint Air Traffic Flow & Capacity Management

DGX agent

arXiv:2608.09315v1 Announce Type: new Abstract: While mathematical models act as vital decision support systems for operational Air Traffic Flow and Capacity Management (ATFCM), existing approaches is

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

ATLAS: Agentic Taxonomy of Large-Scale Software Ecosystems

DGX agent

arXiv:2606.21597v2 Announce Type: replace-cross Abstract: The open-source ecosystem on GitHub lacks a systematic hierarchical taxonomy of software repositories. GitHub Topics, the dominant organizatio

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models

DGX agent

arXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and b

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Automating Deception: Scalable Multi-Turn LLM Jailbreaks

DGX agent

arXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts

DGX agent

arXiv:2601.22758v2 Announce Type: replace Abstract: Large language model agents repeatedly encounter related tasks, yet systems that learn from trajectories commit every lesson to one predefined artif

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks

DGX agent

arXiv:2603.00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduc

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics

DGX agent

arXiv:2608.09638v1 Announce Type: new Abstract: Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reason

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

b10356

DGX agent

ci : target ROCm 7.14 for build and release (#25775) Switch ROCm from 7.2.1 to 7.14 ROCm 7.14 is the first production release using TheRock build system. It can be installed using multi-arch deliverab

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

b10357

DGX agent

opencl: transpose the K tile in local memory for FA prefill kernels (#26428) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED ma

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

b10358

DGX agent

Address review comment of PR 25532 (#26852) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCFramework L

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

b10359

DGX agent

ggml-webgpu: fix CI errors from #25025 and #25262 (#26566) test new flash_attn test rebase and fix to disable subgrou matrices when max_kv_tile == 0 delete log output Add i32 support to cpy and enable

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

b10360

DGX agent

common/peg : suppress incomplete escape sequences (#26780) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iO

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

b10361

DGX agent

model : fix SWA not being enabled for EXAONE 4.5 (#26848) model : fix SWA not being enabled for EXAONE 4.5 load_arch_hparams tests hparams.n_layer() == 64 before LLM_KV_NEXTN_PREDICT_LAYERS has been r

model-releasesllama-cpp-releases
11 Aug 2026
Model Releases

Back to the Future: A workbook time machine for spread sheet creation benchmarks

DGX agent

arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived obj

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

BAG: Budget-Aware Gating for Diffusion Caching

DGX agent

arXiv:2608.09231v1 Announce Type: new Abstract: Diffusion caching is a lightweight strategy that accelerates Diffusion Transformers (DiTs) by reusing intermediate features across denoising steps, but

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

BAP-MOS: Bandit-Based Adaptive Prompting for Boundary-Sensitive Multi-Organ Segmentation

DGX agent

arXiv:2608.08191v1 Announce Type: new Abstract: Multi-organ ultrasound segmentation remains challenging when anatomically adjacent structures must be delineated jointly, as localized boundary errors c

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Bayesian Symbolic Regression with Entropic Reinforcement Learning

DGX agent

arXiv:2608.09617v1 Announce Type: new Abstract: Symbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of inputs. Unlike

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

DGX agent

arXiv:2608.09888v1 Announce Type: cross Abstract: We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference time continuou

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Benchmarking In-context Experiential Learning Through Repeated Product Recommendations

DGX agent

arXiv:2511.22130v2 Announce Type: replace Abstract: To navigate ever-shifting real-world environments, agents must grapple with incomplete knowledge and adapt their strategies through experience. Howe

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms

DGX agent

arXiv:2508.16481v3 Announce Type: replace Abstract: Ensuring the safe use of agentic systems requires a thorough understanding of the range of malicious behaviors these systems may exhibit. In this pa

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Beyond Direct Identifiers: Probabilistic Privacy Risk Estimation for Privacy-Conscious LLM Query Delegation

DGX agent

arXiv:2608.09140v1 Announce Type: cross Abstract: Recent work on protecting privacy during user-LLM interactions often focuses on direct, explicit identifiers: the personally-identifiable information

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

DGX agent

arXiv:2608.09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expecte

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics

DGX agent

arXiv:2512.05098v2 Announce Type: replace-cross Abstract: In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily targe

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents

DGX agent

arXiv:2508.04412v3 Announce Type: replace Abstract: The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serialised user

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Beyond Routing: Decoupling Expert Dispatch and Aggregation in Sparse Mixture-of-Experts

DGX agent

arXiv:2608.08853v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs. We study wheth

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

DGX agent

arXiv:2603.12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge an

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction

DGX agent

arXiv:2608.08459v1 Announce Type: cross Abstract: Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domain

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents

DGX agent

arXiv:2608.09292v1 Announce Type: cross Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. H

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

DGX agent

arXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

DGX agent

arXiv:2608.09555v1 Announce Type: new Abstract: External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Yet their effe

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Biologically Informed Representation Learning for Robust Cross-Center Generalization of MALDI-TOF Mass Spectrometry

DGX agent

arXiv:2608.08182v1 Announce Type: cross Abstract: Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial identificati

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Blumira launches Hearth, an AI command center that spans rival security tools

DGX agent

Security operations platform startup Blumira Inc. today launched Hearth, a vendor-agnostic artificial intelligence command center that lets security teams investigate and act across their existing too

model-releasessiliconangle
11 Aug 2026
Model Releases

Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search

DGX agent

arXiv:2603.24084v2 Announce Type: replace Abstract: Empirical evaluation in multi-objective search (MOS) has historically suffered from fragmentation, relying on heterogeneous problem instances with i

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts

DGX agent

arXiv:2608.09510v1 Announce Type: cross Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and re

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Building Agent Skills and testing them is hard, but it doesn't have to be. Listen to Arjun Patel demo Cultivar, an open source tool develope…

DGX agent

Building Agent Skills and testing them is hard, but it doesn't have to be. Listen to Arjun Patel demo Cultivar, an open source tool developed at Pinecone to help benchmark agent skills in sandboxes. T

model-releasespinecone--x
11 Aug 2026
Model Releases

CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation

DGX agent

arXiv:2608.09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, sup

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?

DGX agent

Hi! I'm really obsessed with LLM hallucinations for the last 6 days 😭 I started by designing system prompts to attack hallucinations but failed, obviously. Now I tried reading logprobs and... I think

model-releasesr-localllama
11 Aug 2026
Model Releases

Can Graph Learning Learn Circuits?

DGX agent

arXiv:2608.08536v1 Announce Type: new Abstract: Circuit localization is a mechanistic interpretability task whose goal is to identify a sparse subgraph of a transformer's computation graph sufficient

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

DGX agent

arXiv:2608.08160v1 Announce Type: cross Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. Howev

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Can Open-Weight Models Compete on Financial Text Comprehension?

DGX agent

arXiv:2608.08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Can We Optimize the Performance-Carbon Emission Break-Even Point?: The Quest for Greener LLMs

DGX agent

arXiv:2608.08744v1 Announce Type: cross Abstract: The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially exceeds th

model-releasesarxiv-cs-ai
11 Aug 2026
← Previous
1…45678…461
Next →