AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “automated”

GridTimelineEvolution
4,979 results
Model Releases

Codex is becoming a productivity tool for everyone

DGX agent

OpenAI's Codex is evolving beyond code generation to become a general productivity tool accessible to non-programmers for knowledge work tasks. The tool leverages large language models to assist with

model-releasesopenai
2 Jun 2026
Agents
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Constitutional Black-Box Monitoring for Scheming in LLM Agents

DGX agent

arXiv:2603.00829v2 Announce Type: replace-cross Abstract: Safe deployment of Large Language Model (LLM) agents in autonomous settings requires reliable oversight mechanisms. A central challenge is det

agentsarxiv-cs-ai
2 Jun 2026
Model Releases

Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem

DGX agent

arXiv:2603.16572v2 Announce Type: replace-cross Abstract: Agent skills extend local AI agents, such as Claude Code and OpenClaw, with additional functionality. Their growing popularity has led to dedi

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

CountGD++: Generalized Prompting for Open-World Counting

DGX agent

arXiv:2512.23351v2 Announce Type: replace Abstract: The flexibility and accuracy of methods for automatically counting objects in images and videos are limited by the way the object can be specified.

agentsarxiv-cs-cv
2 Jun 2026
Model Releases

Cross-Generational Transfer of Adversarial Attacks Reveals Non-Monotonic Safety Alignment in LLMs

DGX agent

arXiv:2606.00813v1 Announce Type: cross Abstract: Safety alignment in LLMs does not improve monotonically across model generations. Studying four generations of Google's Gemma family (7B-31B) with qua

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?

DGX agent

arXiv:2505.16915v3 Announce Type: replace-cross Abstract: While recent Text-to-Image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, they struggle with the lo

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Early Prediction of Liver Cirrhosis Up to Two Years in Advance: A Machine Learning Study Benchmarking Against the FIB-4 and APRI Scores

DGX agent

arXiv:2601.00175v2 Announce Type: replace Abstract: Objective: Develop and evaluate machine learning (ML) models for predicting incident liver cirrhosis (LC) one and two years prior to diagnosis using

model-releasesarxiv-cs-lg
2 Jun 2026
Research

Edge-Based QoS-Aware Adaptive Task Placement: A Closed-Loop Control in Multi-Robot Systems

DGX agent

arXiv:2606.00552v1 Announce Type: cross Abstract: Multi-robot systems (MRS) increasingly offload compute-intensive perception tasks to edge nodes to meet strict time-sensitive Quality-of-Service (QoS)

researcharxiv-cs-ro
2 Jun 2026
Model Releases

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning

DGX agent

arXiv:2604.03893v2 Announce Type: replace Abstract: Current multimodal benchmarks for scientific reasoning primarily evaluate local information extraction -- models recognize symbols and values and th

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

FigSIM: A Dataset for Fine-grained Suicide Severity and Figurative Language in Suicide Memes

DGX agent

arXiv:2606.02523v1 Announce Type: new Abstract: Suicide memes are memes used to express suicide-related thoughts or comment on suicide-related issues. Suicide memes are increasingly common on social m

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

FVSpec: Real-World Property-Based Tests as Lean Challenges

DGX agent

arXiv:2606.01008v1 Announce Type: cross Abstract: We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first scrape 11,039 property-based tes

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Hierarchical Online Prompt Mutation with Dual-Loop Feedback for Guardrailed Evidence Document Generation: A Production-Evaluation Case Study

DGX agent

arXiv:2606.01472v1 Announce Type: cross Abstract: High-stakes production document-generation systems require language models to be adaptive, evidence-grounded, and auditable. We present HOPM, a hierar

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

HLL: Can Agents Cross Humanity's Last Line of Verification?

DGX agent

arXiv:2606.02449v1 Announce Type: new Abstract: Multimodal agents are increasingly expected to operate interfaces on behalf of users, raising a central deployment question: can they truly substitute f

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

How Baz improved its AI Agent Code Review accuracy using Amazon Bedrock AgentCore

DGX agent

This post walks through how Baz built their Spec Review agent using Amazon Bedrock and Amazon Bedrock AgentCore. We'll cover the architecture decisions, implementation details, and the business outcom

agentsaws-ml-blog
2 Jun 2026
Tutorials

Human in the Loop Adaptive Optimization for Improved Time Series Forecasting

DGX agent

arXiv:2505.15354v2 Announce Type: replace Abstract: Time series forecasting models often produce systematic, predictable errors even in critical domains such as energy, finance, and healthcare. We int

tutorialsarxiv-cs-lg
2 Jun 2026
Safety

Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning

DGX agent

arXiv:2606.00334v1 Announce Type: cross Abstract: Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

DGX agent

arXiv:2606.01703v1 Announce Type: cross Abstract: We address the challenge of generating high-fidelity, long-form soundtracks that remain coherent across scene transitions. Existing AI music systems a

model-releasesarxiv-cs-ai
2 Jun 2026
Research

LLM agents patch security bugs, pass all tests, but still leave the vulnerability open [R]

DGX agent

Research demonstrates that LLM-based agents can generate functionally correct patches that pass all tests while still containing security vulnerabilities, challenging the assumption that test-passing

researchr-machinelearning
2 Jun 2026
Model Releases

LLM Consortium for Software Design Refinement: A Controlled Experiment on Multi-Agent Collaboration Topologies

DGX agent

arXiv:2606.01490v1 Announce Type: cross Abstract: We present a controlled experiment evaluating 12 multi-agent LLM collaboration topologies for software architecture design. Using a 2imes2imes2 factor

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Monitoring Agentic Systems Before They're Reliable

DGX agent

arXiv:2606.02494v1 Announce Type: cross Abstract: Agentic systems entering production typically operate as partially integrated assemblies where structural defects, not task-level errors, dominate the

agentsarxiv-cs-ai
2 Jun 2026
Research

Not What, But How: A Communicative Audit of LLM Response Framing

DGX agent

arXiv:2606.02493v1 Announce Type: new Abstract: Large language models (LLMs) are being increasingly used to answer subjective, information-seeking questions, where users are sensitive to how responses

researcharxiv-cs-cl
2 Jun 2026
Model Releases

Order within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery Localization

DGX agent

arXiv:2606.02178v1 Announce Type: cross Abstract: Recent advancements in generative AI have led to image editing models capable of producing realistic forgeries that evade traditional image forgery lo

model-releasesarxiv-cs-ai
2 Jun 2026
Research

Peacemaker at ATE-IT: Automatic term extraction from Italian text for waste management data using encoder model

DGX agent

arXiv:2606.01469v1 Announce Type: new Abstract: The development of automatic term extraction has become increasingly important in modern technology. Automatic term extraction can be found in virtually

researcharxiv-cs-cl
2 Jun 2026
Safety

Perspective on Bias in Biomedical AI: Preventing Downstream Healthcare Disparities

DGX agent

arXiv:2604.14514v2 Announce Type: replace Abstract: Healthcare disparities persist across socioeconomic boundaries, often attributed to unequal access to screening, diagnostics, and therapeutics. Howe

safetyarxiv-cs-ai
2 Jun 2026
Agents

RDA: Reward Design Agent for Reinforcement Learning

DGX agent

arXiv:2606.01672v1 Announce Type: new Abstract: Reinforcement learning has enabled the acquisition of impressive robotic skills, but typically requires hand-crafted reward functions that are slow to d

agentsarxiv-cs-lg
2 Jun 2026
Research

Rethinking Evaluation Paradigms in IBP-based Certified Training

DGX agent

arXiv:2606.02134v1 Announce Type: cross Abstract: Deep neural networks achieve strong performance on many supervised learning tasks but remain vulnerable to adversarial perturbations. Neural network v

researcharxiv-cs-ai
2 Jun 2026
Safety

RL-ACRGNet: Reinforcement Learning-Based Chest Radiology Report Generation Network

DGX agent

arXiv:2606.02035v1 Announce Type: new Abstract: Medical imaging interpretation is a foundational pillar of modern clinical diagnostics, yet the manual generation of radiology reports remains a time-co

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

RoboBenchMart: Benchmarking Robots in Retail Environment

DGX agent

arXiv:2511.10276v2 Announce Type: replace-cross Abstract: Most existing robotic manipulation benchmarks focus on tabletop or household scenarios. While these setups have driven impressive progress, it

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

RocketSmith: An Agentic System for High-Powered Rocket Design and Manufacturing

DGX agent

arXiv:2606.00097v1 Announce Type: new Abstract: This work presents RocketSmith, an agentic system capable of the design, manufacturing, and optimization processes in high powered rocket development. T

agentsarxiv-cs-ro
2 Jun 2026
Model Releases

RubricMiddleware helps your agent verify task completion with a grader subagent This is similar to /goal in claude code or codex, but for de…

DGX agent

RubricMiddleware is a LangChain feature that enables agents to verify task completion by delegating grading to a specialized subagent using predefined rubrics. This approach parallels the goal-verific

model-releasesharrison-chase--x
2 Jun 2026
Model Releases

Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers

DGX agent

arXiv:2606.00902v1 Announce Type: new Abstract: General-purpose VLMs remain unreliable for biomedical research because valid answers in scientific papers depend on evidence split across figures, table

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

Scaling Agentic Capabilities via Grounded Interaction Synthesis

DGX agent

arXiv:2606.02001v1 Announce Type: new Abstract: General agentic intelligence hinges on the ability to interact with diverse real-world tools to complete complex tasks, a capability fundamentally tied

agentsarxiv-cs-cl
2 Jun 2026
Local Ai

scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns

DGX agent

arXiv:2603.17893v2 Announce Type: replace-cross Abstract: Methodology bugs in scientific Python code produce plausible but incorrect results that traditional linters and static analysis tools cannot d

local-aiarxiv-cs-ai
2 Jun 2026
Safety

SHERLOCK: Towards Dynamic Knowledge Adaptation in LLM-enhanced E-commerce Risk Management

DGX agent

arXiv:2510.08948v4 Announce Type: replace-cross Abstract: Effective e-commerce risk management requires in-depth case investigations to identify emerging fraud patterns in highly adversarial environme

safetyarxiv-cs-ai
2 Jun 2026
Model Releases

SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes

DGX agent

arXiv:2606.01912v1 Announce Type: new Abstract: Smart homes are evolving toward complex state-dependent living environments, requiring Large Language Models (LLMs) to reason over user intent, preferen

model-releasesarxiv-cs-ai
2 Jun 2026
Agents

SortingHat: Redefining Operating Systems Education with a Tailored Digital Teaching Assistant

DGX agent

arXiv:2606.00015v1 Announce Type: cross Abstract: Operating Systems (OS) courses are among the most challenging in computer science education due to the complexity of internal structures and the diver

agentsarxiv-cs-ai
2 Jun 2026
Agents

SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale

DGX agent

arXiv:2602.23866v2 Announce Type: replace-cross Abstract: Software engineering agents (SWE) are improving rapidly, with recent gains largely driven by reinforcement learning (RL). However, RL training

agentsarxiv-cs-cl
2 Jun 2026
Model Releases

TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks

DGX agent

arXiv:2606.02384v1 Announce Type: new Abstract: Progress in tabular machine learning has largely focused on increasingly sophisticated model architectures. At the same time, feature engineering remain

model-releasesarxiv-cs-lg
2 Jun 2026
Applications

TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination

DGX agent

arXiv:2606.01737v1 Announce Type: new Abstract: Traffic accident liability analysis is a critical yet challenging task in intelligent transportation and legal assistance. Existing methods often suffer

applicationsarxiv-cs-ai
2 Jun 2026
Tutorials

Travelers deploys AI-powered claims countrywide with OpenAI

DGX agent

Travelers Insurance has deployed an AI system powered by OpenAI technology across its operations nationwide to streamline and improve claims processing. The implementation leverages OpenAI's capabilit

tutorialsopenai
2 Jun 2026
Model Releases

WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis

DGX agent

arXiv:2606.01869v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural lan

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

A Novel Global Context-aware Deep Neural Network for Enhanced Brain Tumor Segmentation using Magnetic Resonance Images

DGX agent

arXiv:2605.30510v1 Announce Type: cross Abstract: Brain cancer's severity necessitates precise brain tumor segmentation, which is crucial for effective brain tumor diagnosis. Manual identification, bu

model-releasesarxiv-cs-ai
1 Jun 2026
Research

A Padding Method for Enhanced Encoding of Inorganic Structures with Varying Chemical Compositions

DGX agent

arXiv:2605.30743v1 Announce Type: cross Abstract: Designing novel inorganic materials through generative models remains an important challenge for material science, driven by the complexity and divers

researcharxiv-cs-cl
1 Jun 2026
Safety

AI Loss of Control Incident Management: Response & Resilience

DGX agent

arXiv:2605.30406v1 Announce Type: cross Abstract: Recent research demonstrating AI systems exhibiting deception and shutdown resistance suggests that AI loss of control (LOC) is an urgent policy conce

safetyarxiv-cs-ai
1 Jun 2026
Applications

Astra: a generalizable report generation foundation model for 3D computed tomography

DGX agent

arXiv:2605.31437v1 Announce Type: new Abstract: CT interpretation requires radiologists to review hundreds of volumetric slices per examination, making reporting time-consuming and highly expertise-de

applicationsarxiv-cs-cv
1 Jun 2026
Research

Attention-based optimizer for symmetry finding

DGX agent

arXiv:2605.30429v1 Announce Type: cross Abstract: Finding symmetries is crucial for understanding physical models. In this work, we present an optimization framework that searches Pauli symmetries of

researcharxiv-cs-lg
1 Jun 2026
Local Ai

Beyond Accuracy: Evaluating Efficiency, Robustness and Explainability in Deep Learning for Malaria Diagnosis

DGX agent

arXiv:2605.30734v1 Announce Type: cross Abstract: Malaria remains a leading cause of mortality in sub-Saharan Africa, where scarce diagnostic infrastructure makes timely, accurate diagnosis particular

local-aiarxiv-cs-cv
1 Jun 2026
Safety

Biases in the Blind Spot: Detecting What LLMs Fail to Mention

DGX agent

arXiv:2602.10117v5 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often provide chain-of-thought (CoT) reasoning traces that appear plausible, but may hide internal biases. We cal

safetyarxiv-cs-ai
1 Jun 2026
← Previous
1…7677787980…104
Next →