AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,566 results
2 Jul 2026

Wordle 1,839 4/6 ⬛⬛🟨⬛🟨 ⬛⬛🟨⬛⬛ ⬛🟨⬛🟨🟩 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

I cannot provide a meaningful summary for this entry as the content appears to be a personal Wordle game result (puzzle #1,839 solved in 4 attempts) rather than substantive knowledge base material. Th

WorkBench Revisited: Workplace Agents Two Years On

Model ReleasesDGX agent

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to

XSkill: Continual Learning from Experience and Skills in Multimodal Agents

Model ReleasesDGX agent

arXiv:2603.12056v3 Announce Type: replace Abstract: Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible orchestr


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

Model ReleasesDGX agent

arXiv:2607.00664v1 Announce Type: new Abstract: We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese

You really need your own benchmarks. If you are translating hieroglyphics, use Gemini 3.5 Flash. If you are running a vending machine use Op…

Model ReleasesDGX agent

You really need your own benchmarks. If you are translating hieroglyphics, use Gemini 3.5 Flash. If you are running a vending machine use Opus 4.8. (This is one reason why I am skeptical of just swapp

Your coding agent bill doubled and nobody can tell you why. Here's the actual reason: Claude Code, Cursor, and Copilot all log activity in d…

Model ReleasesDGX agent

Your coding agent bill doubled and nobody can tell you why. Here's the actual reason: Claude Code, Cursor, and Copilot all log activity in different formats. The second your team uses more than one (t

Z.ai launches ZCode, an 'Agentic Development Environment' optimized for its new GLM-5.2 model; Z.ai's GLM Coding Plan costs from 16.20 to 144 per month (Michael Nuñez/VentureBeat)

Model ReleasesDGX agent

Michael Nuñez / VentureBeat: Z.ai launches ZCode, an “Agentic Development Environment” optimized for its new GLM-5.2 model; Z.ai's GLM Coding Plan costs from 16.20 to 144 per month — The move marks th

ZO-Act: Efficient Zeroth-Order Fine-Tuning via One-Shot Activation-Informed Low-Rank Subspaces

Model ReleasesDGX agent

arXiv:2607.01125v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables fine-tuning large language models when backpropagation is unavailable or memory-prohibitive, but existing methods

1 Jul 2026

A Large-Language-Model Supported Personalized Driving Framework for Lane Change in Highway Scenarios

Model ReleasesDGX agent

arXiv:2606.31483v1 Announce Type: new Abstract: Personalized driving can improve the user acceptance of automated driving systems. However, existing methods still provide limited support for translati

A Realistic Protocol for Evaluation of Weakly Supervised Object Localization

Model ReleasesDGX agent

arXiv:2404.10034v3 Announce Type: replace Abstract: Weakly Supervised Object Localization (WSOL) allows training deep learning models for classification and localization (LOC) using only global class-

A Reproducible Benchmark of Lightweight CNNs: Accuracy, Efficiency, and the Impact of Pretrained Initialization

Model ReleasesDGX agent

arXiv:2505.03303v3 Announce Type: replace-cross Abstract: Lightweight convolutional neural networks are often compared using results obtained with different training recipes, input settings, and pretr

A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols

Model ReleasesDGX agent

arXiv:2606.31763v1 Announce Type: new Abstract: Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constraints and experi

A Semantic-Layer-Mediated Agent for Natural Language to SQL over Heterogeneous Enterprise Databases

Model ReleasesDGX agent

arXiv:2606.31041v1 Announce Type: new Abstract: Natural language-to-SQL (NL2SQL) over real-world enterprise databases remains significantly more challenging than on academic benchmarks. Enterprise sch

A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection

Model ReleasesDGX agent

arXiv:2606.30837v1 Announce Type: cross Abstract: The number of trees is a central computational parameter in Random Forests: increasing it reduces finite-ensemble variability but increases training a

A swap-adversarial framework for improving domain generalization in electrocorticography-based Parkinson's disease classification

Model ReleasesDGX agent

arXiv:2602.10528v2 Announce Type: replace-cross Abstract: We propose a novel swap-adversarial framework that mitigates high inter-subject variability and the high-dimensional low-sample-size problem i

A Systematic Approach to Multi-Agent AI from Advanced Regulatory Control Theory: Safe and Auditable LLM Operator Agents for Process Control

Model ReleasesDGX agent

arXiv:2606.30877v1 Announce Type: cross Abstract: Recent literature shows that large language models (LLMs) are useful for general-purpose tasks yet perform poorly on specific domain ones. One reason

A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management

Model ReleasesDGX agent

arXiv:2606.30997v1 Announce Type: new Abstract: We present a three-phase deep reinforcement learning system for personalized portfolio management that addresses three limitations shared by all prior f

A time-series classification framework for individual-level absenteeism prediction under severe class imbalance

Model ReleasesDGX agent

arXiv:2606.31532v1 Announce Type: new Abstract: Staff absenteeism imposes substantial operational costs in high-demand work environments such as healthcare, emergency services, meat processing, constr

A Transferable Learned Temporal Prior for Transmission Reconstruction and Decision-Relevant Uncertainty in Real Outbreak Labels

Model ReleasesDGX agent

arXiv:2606.30842v1 Announce Type: new Abstract: Outbreak transmission reconstruction treats epidemiological timing and transmission labels as deterministic ground truth; neither has been systematicall

Absorption-Feature-Guided Distance-Decoupled Estimation and Band Selection for LWIR Hyperspectral Passive Ranging

Model ReleasesDGX agent

arXiv:2606.31824v1 Announce Type: new Abstract: Long-wave infrared (LWIR) hyperspectral observations contain distance-dependent atmospheric absorption signatures, providing a physical basis for long-r

Accelerometry-Derived Digital Biomarkers for Cardiometabolic Risk: A Population-Representative Tabular Benchmark with Uncertainty Quantification

Model ReleasesDGX agent

arXiv:2606.30702v1 Announce Type: cross Abstract: Structured tabular data dominates clinical medicine, yet existing benchmarks fail to reflect real-world properties like complex survey sampling, demog

Adaptive Cluster-First Route-Second Decomposition for Industrial-Scale Vehicle Routing

Model ReleasesDGX agent

arXiv:2606.31820v1 Announce Type: new Abstract: Large-scale capacitated vehicle routing problems (CVRPs) are commonly addressed using cluster-first route-second (CFRS) approaches that split a routing

Again, 2 weeks missed. No rehab. Rib injury, which really affects ability to rotate (comfortably) and he's just raking. He's on a heater.

Model ReleasesDGX agent

Again, 2 weeks missed. No rehab. Rib injury, which really affects ability to rotate (comfortably) and he's just raking. He's on a heater. It is kind of unfathomable that Chase DeLauter missed two week

AgentBound: Verifiable Behavioral Governance for Autonomous AI Agents

Model ReleasesDGX agent

arXiv:2606.30970v1 Announce Type: new Abstract: Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external communications

“Agentic kernel optimization is the future of on-device inference” @xenovacom used Fable 5 to write kernels that pushed Gemma 4 to a massive…

Model ReleasesDGX agent

“Agentic kernel optimization is the future of on-device inference” @xenovacom used Fable 5 to write kernels that pushed Gemma 4 to a massive 255 tok/s on WebGPU with M4. He shared the demo, so you can

Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping

Model ReleasesDGX agent

arXiv:2606.31200v1 Announce Type: new Abstract: Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based me

Agentic retrieval is changing the way retrieval-augmented applications are built, especially in domains like legal and fintech, where agents…

Model ReleasesDGX agent

Agentic retrieval is changing the way retrieval-augmented applications are built, especially in domains like legal and fintech, where agents need to autonomously navigate large, evolving knowledge bas

AION: Aerial Indoor Object-Goal Navigation Using Dual-Policy Reinforcement Learning

Model ReleasesDGX agent

arXiv:2601.15614v3 Announce Type: replace Abstract: Object-Goal Navigation (ObjectNav) requires an agent to autonomously explore an unknown environment and navigate toward target objects specified by

AlloyDB AI Functions - now with revolutionary performance boosts and cost savings

Model ReleasesDGX agent

AlloyDB is an AI-native database—it isn’t just a passive data store, it intelligently understands and processes your data. With AlloyDB, you get industry-leading vector and hybrid search, near 100% ac

An Empirical Study of Security Calibration in Large Language Models for Code

Model ReleasesDGX agent

arXiv:2606.31159v1 Announce Type: cross Abstract: Large Language Models (LLMs) are rapidly transforming software development, yet their use in security-critical contexts raises a key question: do mode

An Executable Benchmarking Suite for Tool-Using Agents

Model ReleasesDGX agent

arXiv:2605.11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often con

Anthropic launches Claude Sonnet 5 AI model with coding, safety upgrades as Fable and Mythos controls lifted

Model ReleasesDGX agent

Anthropic PBC today debuted Claude Sonnet 5, a midrange large language model that outperforms its predecessor in several areas. The LLM will be the default option in the consumer tiers of the company’

Anthropic says Fable 5 will be available via usage credits for Claude users from July 7, and is working with partners to draft an AI jailbreak severity standard (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic says Fable 5 will be available via usage credits for Claude users from July 7, and is working with partners to draft an AI jailbreak severity standard — On Friday, June 12, the US

Anthropic says it is rolling back a covert Claude Code tracking feature to identify users based in China or affiliated with Chinese AI labs, after backlash (Juro Osawa/The Information)

Model ReleasesDGX agent

Juro Osawa / The Information: Anthropic says it is rolling back a covert Claude Code tracking feature to identify users based in China or affiliated with Chinese AI labs, after backlash — Anthropic is

Anthropic says 'some routine tasks like coding and debugging' on Fable 5 'will fall back to Opus 4.8' in 'the near term' as it works to 'reduce false positives' (@anthropicai)

Model ReleasesDGX agent

@anthropicai: Anthropic says “some routine tasks like coding and debugging” on Fable 5 “will fall back to Opus 4.8” in “the near term” as it works to “reduce false positives” — Claude Fable 5 will be

ARC-AGI-3 is built different, it has dumbfounded almost all regular attempts so far because it's so much harder than anything that came befo…

Model ReleasesDGX agent

ARC-AGI-3 is built different, it has dumbfounded almost all regular attempts so far because it's so much harder than anything that came before. It has no rules, it's agentic and has no explicit goals,

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

Model ReleasesDGX agent

arXiv:2606.31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) model

Armadin details full sandbox escape in Claude Cowork but Anthropic disputes risk

Model ReleasesDGX agent

Security researchers at Armadin Inc. today detailed an attack chain that runs arbitrary commands as root inside the sandbox behind Anthropic PBC’s Claude Cowork, escaping the isolation layer, with a s

Artificial Intelligence in Sports: Insights from a Quantitative Survey among Sports Students in Germany about their Perceptions, Expectations, and Concerns regarding the Use of AI Tools

Model ReleasesDGX agent

arXiv:2503.05785v2 Announce Type: replace-cross Abstract: Generative Artificial Intelligence (AI) tools such as ChatGPT, Copilot, or Gemini have a crucial impact on academic research and teaching. Emp

As generative AI tools continue to evolve, we believe it's more important than ever to know what's AI-generated and what isn't. That’s why @…

Model ReleasesDGX agent

As generative AI tools continue to evolve, we believe it's more important than ever to know what's AI-generated and what isn't. That’s why @GoogleDeepMind launched SynthID in 2023—a technology that ad

AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

Model ReleasesDGX agent

arXiv:2606.31551v1 Announce Type: new Abstract: Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents become increasingly capable at software

AxDafny: Agentic Verified Code Generation in Dafny

Model ReleasesDGX agent

arXiv:2606.32007v1 Announce Type: new Abstract: We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny

BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation

Model ReleasesDGX agent

arXiv:2606.30850v1 Announce Type: new Abstract: Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce epistemic unce

Benchmarking Large Language Models on Floating-Point Error Classification

Model ReleasesDGX agent

arXiv:2606.31308v1 Announce Type: new Abstract: This paper investigates the capability of Large Language Models (LLMs) to detect and classify floating-point errors statically in software code. We intr

Best Claude use case ever: learning to use Microsoft Teams for first time 🙃🤣 - from @_catwu at AIE w/ @swyx & @trq212

Model ReleasesDGX agent

This post humorously describes using Claude AI as a learning tool to understand Microsoft Teams functionality for first-time users, shared during an AI Engineer event discussion. The post suggests Cla

Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

Model ReleasesDGX agent

arXiv:2606.31338v1 Announce Type: cross Abstract: Recent music audio-language models achieve high accuracy on instrument question-answering benchmarks, but it remains unclear whether this reflects rob

Beyond Clean Text: Evaluating Encoder and Decoder Robustness for Bangla Event Detection in Noisy Text

Model ReleasesDGX agent

arXiv:2606.30914v1 Announce Type: new Abstract: Event detection (ED) systems are typically evaluated on clean, curated text, leaving their robustness to real-world noise largely unexplored, particular

Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization

Model ReleasesDGX agent

arXiv:2606.31002v1 Announce Type: new Abstract: Theorem-proving benchmarks evaluate proof search against fixed formal statements, but natural-language-to-Lean formalization must generate the formal st

Beyond expert users: agents should help users construct preferences, not just elicit them

Model ReleasesDGX agent

arXiv:2606.30863v1 Announce Type: new Abstract: Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task

Beyond Single Character: Evaluating MLLMs for Sentence-Level Oracle Bone Inscription Understanding

Model ReleasesDGX agent

arXiv:2606.31169v1 Announce Type: new Abstract: Existing AI-assisted oracle bone inscription (OBI) visual recognition and understanding studies mainly focus on character-level, ignoring the long-form

Beyond Static Prompts: Building Scale-Proof, Polymorphic Multi-Agent Systems with Google's ADK

Model ReleasesDGX agent

As enterprise generative AI transitions from simple, conversational chatbots to autonomous multi-agent workflows, developers face a critical bottleneck: scale. In a production environment, an enterpri

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

Model ReleasesDGX agent

arXiv:2606.22723v2 Announce Type: replace Abstract: Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended,

Bridging Scientific Heritage: An Arabic--Russian Parallel Corpus and LLM Benchmark for Sustainable Knowledge Transfer

Model ReleasesDGX agent

arXiv:2606.30943v1 Announce Type: new Abstract: Russian and Arabic are among the major languages of scientific communication. Language barriers impede the exchange of research results between these co

Building an ASR Solution for Training and Assessing Children's Reading

Model ReleasesDGX agent

arXiv:2606.31508v1 Announce Type: new Abstract: Automatic speech recognition for children's reading remains underdeveloped for most African languages, including Bambara, despite its potential value fo

Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?

Model ReleasesDGX agent

arXiv:2606.31371v1 Announce Type: cross Abstract: When large language model (LLM) agents adapt their behavior through evaluator feedback, systematic evaluator biases propagate into the agent's learned

Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models

Model ReleasesDGX agent

arXiv:2606.31630v1 Announce Type: new Abstract: Language models increasingly write probabilistic programs (in NumPyro, Stan, or Pyro), but a program that compiles, runs, and passes every unit test can

Can Physician Expertise Improve Machine Learning Identification of Delirium?

Model ReleasesDGX agent

arXiv:2606.30651v1 Announce Type: cross Abstract: Delirium is common in hospitalized patients and is often missed in routine care. We present a user-centered interactive machine learning (UC-iML) fram

@_catwu @simonw and I will be doing a fireside chat about 'This year in Claude' from 12:30pm-1:30pm at AIE in Expo Stage 2. We'll be coverin…

Model ReleasesDGX agent

@_catwu @simonw and I will be doing a fireside chat about 'This year in Claude' from 12:30pm-1:30pm at AIE in Expo Stage 2. We'll be covering a really wide range of topics and I think it will be reall

CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes

Model ReleasesDGX agent

arXiv:2606.31435v1 Announce Type: new Abstract: Data refinement involves executing multi-step recipes over evolving text states, where both composition and execution order of processing operators dete

CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield

Model ReleasesDGX agent

arXiv:2606.31796v1 Announce Type: cross Abstract: We study three complementary techniques for training compute-efficient language models. (1) Selective supervision and per-token efficiency. Selective

← Previous
1…111112113114115…377
Next →