AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,429
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,935
  • Model Releases23,918
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,429
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,935
  • Model Releases23,918
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

88,429Total entries
1Added by human
88,428Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,648 results
6 Jul 2026

'Ghost memory' is a real problem with agents. You might have seen the issue where a long-running agent still confidently repeats a user fact…

Model ReleasesDGX agent

'Ghost memory' is a real problem with agents. You might have seen the issue where a long-running agent still confidently repeats a user fact that stopped being true weeks ago? New research names the f

5 Jul 2026

Fable prompting tips, based on all the guides Anthropic and its employees have given us. Three things you might want to do: 1) review your c…

Model ReleasesDGX agent

Fable prompting tips, based on all the guides Anthropic and its employees have given us. Three things you might want to do: 1) review your claude.md file against prompting best practices so you don't

my contribution to this weekend’s family gathering was using claude to build a family tree that connects everybody the only way to do this w…

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

Yohei Nakajima used Claude AI to construct a comprehensive family tree that connected all family members for a weekend gathering. The post highlights a practical application of AI in organizing and vi

3 Jul 2026

A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory

Model ReleasesDGX agent

arXiv:2607.01935v1 Announce Type: new Abstract: Long term memory lets LLM agents act as persistent assistants, but user facts change. A useful memory system must know what is true now, what used to be

Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent

Model ReleasesDGX agent

arXiv:2602.03001v2 Announce Type: replace-cross Abstract: To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, rely

AgenticDataBench: A Comprehensive Benchmark for Data Agents

Model ReleasesDGX agent

arXiv:2607.01647v1 Announce Type: cross Abstract: Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern so

An Additive MLP-GNN Framework for Characterizing Chemical and Structural Contributions to Aqueous Solubility

ResearchDGX agent

arXiv:2607.02212v1 Announce Type: cross Abstract: Aqueous solubility is a key property in early-stage drug discovery, but most predictive models merge physicochemical descriptors and molecular graph i

Anthropic wants to develop its own drugs

Model ReleasesDGX agent

At the event 'The Briefing: AI for Science' earlier this week, Anthropic announced Claude Science, a new 'AI workbench for scientists' that pulls fragmented tools and datasets into one environment, an

Bringing Agentic Search to Earth Observation Data Discovery

Model ReleasesDGX agent

arXiv:2607.02387v1 Announce Type: cross Abstract: NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding

Controllable Sim Agents with Behavior Latents

Model ReleasesDGX agent

arXiv:2607.02496v1 Announce Type: cross Abstract: Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enabl

Dendritic In-Context Learning in a Single-Layer Spiking Neural Network

Model ReleasesDGX agent

arXiv:2607.02283v1 Announce Type: cross Abstract: In-context learning (ICL) operates via implicit gradient descent embedded in the forward pass of modern AI architectures -- Transformers, Mamba, state

eCream-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

Model ReleasesDGX agent

arXiv:2606.12569v2 Announce Type: replace-cross Abstract: We present eCream-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals. Th

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Model ReleasesDGX agent

arXiv:2607.02440v1 Announce Type: new Abstract: Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a

Extreme Adaptive Transformer for Time Series Forecasting

ApplicationsDGX agent

arXiv:2607.02437v1 Announce Type: new Abstract: Time series forecasting remains challenging when the underlying data contain rare but critical extreme events. This issue is particularly important in h

Geometric Signatures of Reasoning: A Spectral Perspective on Task Hardness

ResearchDGX agent

arXiv:2607.01571v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning enables large language models (LLMs) to solve complex problems by generating intermediate reasoning steps. While much a

Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting

AgentsDGX agent

arXiv:2607.01457v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly applied to resume optimization for applicant tracking systems, introducing hallucination failures distin

Influence of Radial Basis Activation Functions on Intelligent Controller for Robotic Manipulators

Model ReleasesDGX agent

arXiv:2607.02167v1 Announce Type: cross Abstract: This paper presents an intelligent control framework for trajectory tracking of robotic manipulators using radial basis function (RBF) neural networks

MKGR: Multimodal Knowledge-Graph Representation Learning for Cold-Start Protein-Protein Interaction Prediction

Model ReleasesDGX agent

arXiv:2607.01627v1 Announce Type: cross Abstract: Accurate protein-protein interaction (PPI) prediction is central to functional genomics, disease mechanism discovery, and drug development. A difficul

Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation

SafetyDGX agent

arXiv:2607.02460v1 Announce Type: cross Abstract: Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in s

OmniGAIA: Towards Native Omni-Modal AI Agents

Model ReleasesDGX agent

arXiv:2602.22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to i

PPTArena: A Benchmark for PowerPoint Editing

Model ReleasesDGX agent

arXiv:2512.03042v3 Announce Type: replace-cross Abstract: We introduce PPTArena, a benchmark for PowerPoint editing that evaluates how agents modify real slides from natural-language instructions. Unl

Prompt Framing Distorts Count-Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

Model ReleasesDGX agent

arXiv:2607.01240v1 Announce Type: cross Abstract: Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding i

Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness

Model ReleasesDGX agent

arXiv:2510.04484v2 Announce Type: replace-cross Abstract: The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interactio

QTALE: Quantization-Robust Token-Adaptive Layer Execution for LLMs

ResearchDGX agent

arXiv:2602.10431v4 Announce Type: replace Abstract: Large language models (LLMs) demand substantial computational and memory resources, posing challenges for efficient deployment. Two complementary ap

RadiomicNet: A Hybrid Radiomics-Guided Lightweight Architecture for Interpretable Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2607.02185v1 Announce Type: cross Abstract: Deep learning has achieved remarkable performance in medical image segmentation, yet it suffers from critical limitations: mathematical intractability

Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study

Local AiDGX agent

arXiv:2607.02436v1 Announce Type: cross Abstract: Agentic coding assistants are increasingly given extra capabilities, such as browser based testing tools and design oriented system prompts, on the as

Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas

Model ReleasesDGX agent

arXiv:2607.02504v1 Announce Type: cross Abstract: Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on extbf{sp

RedCoder: Automated Multi-Turn Red Teaming for Code LLMs

SafetyDGX agent

arXiv:2507.22063v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) for code generation (i.e., Code LLMs) have demonstrated impressive capabilities in AI-assisted software developme

RusFinChain: A Russian Benchmark for Verifiable Chain-of-Thought Reasoning in Finance with Fuzzy-Aligned Evaluation

Model ReleasesDGX agent

arXiv:2607.01388v1 Announce Type: new Abstract: Multi-step symbolic reasoning is essential for robust financial analysis, yet most benchmarks neglect intermediate reasoning steps. FINCHAIN introduced

Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification

Model ReleasesDGX agent

arXiv:2607.01793v1 Announce Type: new Abstract: LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testin

Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving

SafetyDGX agent

arXiv:2604.03497v2 Announce Type: replace-cross Abstract: Vision-language-model (VLM)-guided reinforcement learning (RL) has recently attracted significant attention for it, replacing brittle hand-cra

SINA: A Fully Automated Circuit Schematic Image to Netlist Generator Using Artificial Intelligence

ResearchDGX agent

arXiv:2607.01609v1 Announce Type: new Abstract: Recent advances in Artificial Intelligence (AI) have revolutionized Electronic Design Automation (EDA), particularly through Large Language Models (LLMs

Spatial Support Matters: Geometry-Aware Graph Fusion for Rainfall Field Reconstruction

ResearchDGX agent

arXiv:2607.01621v1 Announce Type: new Abstract: Fine-scale rainfall reconstruction is critical for urban flood modeling, but real rainfall sensing systems observe the field through incompatible spatia

The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits

Model ReleasesDGX agent

arXiv:2607.02201v1 Announce Type: cross Abstract: The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains fragmented

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue

Model ReleasesDGX agent

arXiv:2607.01345v1 Announce Type: cross Abstract: Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluations often re

we distilled 2.3M Claude Fable 5 reasoning traces into Qwen3-4B - 100% self-consistency @ 512 samples - 0.00 bits output entropy - zero hall…

Model ReleasesDGX agent

we distilled 2.3M Claude Fable 5 reasoning traces into Qwen3-4B - 100% self-consistency @ 512 samples - 0.00 bits output entropy - zero hallucination variance turns out the student is not bounded by t

2 Jul 2026

Adaptive Perturbation Selection for Contrastive Audio Decoding

ResearchDGX agent

arXiv:2607.00247v1 Announce Type: cross Abstract: Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors. While contrastive decoding (CD) offer

AFFMAE: Scalable Vision Pre-Training for High-Resolution Microscopy Segmentation on Desktop Hardware

Model ReleasesDGX agent

arXiv:2602.16249v2 Announce Type: replace Abstract: Self-supervised pretraining has transformed computer vision by enabling data-efficient fine-tuning, yet high-resolution pretraining typically requir

Amortized Maximum Inner Product Search with Learned Support Functions

Model ReleasesDGX agent

arXiv:2603.08001v3 Announce Type: replace Abstract: Maximum inner product search (MIPS) is a crucial subroutine in machine learning, requiring the identification of a vector taken within a database (t

Autonomous Scientific Discovery via Iterative Meta-Reflection

Model ReleasesDGX agent

arXiv:2607.01131v1 Announce Type: cross Abstract: Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and validation.

Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents

Model ReleasesDGX agent

arXiv:2607.00895v1 Announce Type: new Abstract: Hallucination detection for retrieval-augmented generation (RAG) is usually evaluated on natural-language document evidence. However, grounded generatio

Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated Moderation Traces

SafetyDGX agent

arXiv:2607.00481v1 Announce Type: cross Abstract: Jailbreak attacks remain a critical threat to the safe deployment of large language models (LLMs). While prior work has primarily studied attacks and

Cross4D-JEPA: Dense Cross-modal Correspondence Distillation for 4D Point Cloud Representation Learning

ResearchDGX agent

arXiv:2607.00514v1 Announce Type: cross Abstract: Automatic understanding of dynamic 4D point clouds, the 3D-point sequences captured over time by depth sensors and LiDAR, is central to robotics and e

Detecting the Undetectable: Enhancing Unsupervised time series Anomaly Detection via Active Learning

TutorialsDGX agent

arXiv:2607.00720v1 Announce Type: cross Abstract: Despite the increasing sophistication of industrial AI systems, the ability to reliably detect subtle and noisy anomalies in complex time series data

Diffusion-GR2: Diffusion Generative Reasoning Re-ranker

SafetyDGX agent

arXiv:2607.01170v1 Announce Type: cross Abstract: Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list, but they ar

Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shifts

SafetyDGX agent

arXiv:2607.00666v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models often fail to perform the same learned tasks under environmental shifts, such as changes in camera pose and shifts

EVOTS: Evolutionary Transformer Search for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2607.00154v1 Announce Type: cross Abstract: Evolutionary neural architecture design for multivariate time-series forecasting remains underexplored, with most approaches relying on fixed Transfor

FedIA: Importance-Aware Aggregation for Domain-Robust Federated Graph Learning

Model ReleasesDGX agent

arXiv:2509.18171v4 Announce Type: replace Abstract: Federated graph learning (FGL) is a natural paradigm for social-media user graphs, where language communities, regional markets, and service boundar

FRAME: Learning the Adaptation Domain with a Mixture of Fractional-Fourier Experts

Model ReleasesDGX agent

arXiv:2607.00162v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) reparameterizes weight updates in a fixed basis: low-rank adapters operate in the spatial domain, while a recent

From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives

AgentsDGX agent

arXiv:2607.00918v1 Announce Type: cross Abstract: Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consistency and co

FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal

Model ReleasesDGX agent

arXiv:2603.19036v2 Announce Type: replace Abstract: Single image reflection removal (SIRR) is challenging in real scenes, where reflection strength varies spatially and reflection patterns are tightly

GAIA: Geometry-Adaptive Operator Learning for Forward and Inverse Problems

Model ReleasesDGX agent

arXiv:2607.01128v1 Announce Type: new Abstract: Operator learning for partial differential equations (PDEs) on arbitrary geometries builds fast neural surrogates for large-scale simulation. Although r

I just left the final day of the @aiDotEngineer World's Fair Conference in San Francisco. Kudos to @swyx for putting together a world-class …

Model ReleasesDGX agent

I just left the final day of the @aiDotEngineer World's Fair Conference in San Francisco. Kudos to @swyx for putting together a world-class lineup of speakers and workshops! It really was an invigorat

llm-coding-agent 0.1a0

Model ReleasesDGX agent

Release: llm-coding-agent 0.1a0 Another Fable 5 experiment. Now that my LLM library has evolved into more of an agent framework it's time to see what a simple coding agent would look like built on it.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs

ResearchDGX agent

arXiv:2602.05275v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have shown immense promise in universal multimodal retrieval, which aims to find relevant items of various

Multi-Label Node Classification with Label Influence Propagation

Model ReleasesDGX agent

arXiv:2607.00671v1 Announce Type: cross Abstract: Graphs are a complex and versatile data structure used across various domains, with possibly multi-label nodes playing a particularly crucial role. Ex

Not your weights not your GF

IndustryDGX agent

This post likely discusses the concept of model weights and their ownership or control in machine learning contexts, possibly drawing a playful analogy to relationship dynamics. The title suggests com

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping

SafetyDGX agent

arXiv:2607.00881v1 Announce Type: new Abstract: Spatial intelligence remains a persistent challenge for Multimodal Large Language Models (MLLMs), as it requires coherent spatial scene representations

Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning

Local AiDGX agent

arXiv:2607.01191v1 Announce Type: new Abstract: Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resoluti

Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions

ResearchDGX agent

arXiv:2607.00937v1 Announce Type: new Abstract: Persona-driven generations (PDGs) have seen prolific use in research and industry applications, where a large language model (LLM) takes on a 'persona'

← Previous
1…524525526527528…1061
Next →