AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “automated”

GridTimelineEvolution
4,980 results
3 Jun 2026

Diagnosing Knowledge Gaps in LLM Tool Use: An Agentic Benchmark for Novel API Acquisition

Model ReleasesDGX agent

arXiv:2606.03657v1 Announce Type: new Abstract: Large language models for code generation often need to use APIs that are absent from their pretraining data. This requires more than recalling a functi

Enhancing Operational Safety via Agentic Dialogue Hazard Identification Analysis

SafetyDGX agent

arXiv:2606.03812v1 Announce Type: new Abstract: Operational safety in high-stakes domains such as industrial process control, autonomous, and safety-critical systems, demand reliable hazard identifica

EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.02971v1 Announce Type: new Abstract: Extracting reporting obligations from EU legislation is critical for assessing and reducing regulatory reporting burden. However, distinguishing reporti

EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving via Self-Improving LLM Agents

Model ReleasesDGX agent

arXiv:2606.03678v1 Announce Type: new Abstract: Generating safety-critical scenarios is essential for validating and improving autonomous driving systems, yet it inherently requires maximizing adversa

FORGE: Multi-Agent Graduated Exploitation and Detection Engineering

AgentsDGX agent

arXiv:2606.03453v1 Announce Type: cross Abstract: Vulnerability disclosure volumes now far exceed organizational assessment capacity, yet three adjacent research communities (proof-of-concept generati

GN0: Toward a Unified Paradigm for Generation, Evaluation, and Policy Learning in Visual-Language Navigation

Model ReleasesDGX agent

arXiv:2606.03682v1 Announce Type: new Abstract: Embodied navigation connects intelligent agents with the physical world and is fundamental for general robotic intelligence. Limited availability and qu

GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph Theory

Model ReleasesDGX agent

arXiv:2606.03144v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning as

High-Precision APT Malware Attribution with Out-of-Scope Resilience

ResearchDGX agent

arXiv:2606.03523v1 Announce Type: cross Abstract: Early attribution of Advanced Persistent Threat (APT) activity can help defenders prioritise investigation, select countermeasures, and reduce the imp

How Many Trees in a Random Forest? A Revisited Approach with Plateau Search and Optuna Integration

Model ReleasesDGX agent

arXiv:2606.03549v1 Announce Type: new Abstract: Hyperparameter optimization (HPO) for Random Forest faces a specific difficulty in tuning the number of trees: the predictive score typically improves m

Lean-GAP: A Dataset of Formalized Graduate Algebra Problems

ResearchDGX agent

arXiv:2606.02588v1 Announce Type: cross Abstract: We present Lean-GAP (Lean-Graduate Agebra Problems), 430 formalized graduate-level algebra problems from the textbook Abstract Algebra by Dummit and F

LEAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

Model ReleasesDGX agent

arXiv:2606.03303v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit strong informal mathematical reasoning but struggle to generate mechanically verifiable proofs in formal languages

Nanocoder 1.27.0 - skills, daemon + more 🔥

Local AiDGX agent

Nanocoder 1.27.0 is an agentic coding tool available in your terminal that runs on any AI model you choose, whether local models via Ollama or cloud providers like OpenAI and Anthropic. This release i

NeurIPS used uncalibrated AI detector for desk rejections [D]

ResearchDGX agent

NeurIPS 2026 used an AI detector to identify policy violations, resulting in 178 desk-rejected submissions (18.4% of all submissions) and 123 flagged for further review (12.7%) . The detection approac

Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance

SafetyDGX agent

arXiv:2511.10055v2 Announce Type: replace Abstract: The performance of image generation has been significantly improved in recent years. However, the study of image screening is rare, and its performa

Safety Measurements for Fine-tuned LLMs Should be Grounded in Capability

Model ReleasesDGX agent

arXiv:2606.03648v1 Announce Type: cross Abstract: Adapting foundation large language models to a user's task or preferred style through fine-tuning can result in compromising the model's safety. Previ

SAIL: Sound Abstract Interpreters with LLMs

TutorialsDGX agent

arXiv:2511.13663v2 Announce Type: replace-cross Abstract: How to construct globally sound abstract interpreters to safely approximate program behaviors remains a bottleneck in abstract interpretation.

SLU-2K: A Question-Based Benchmark for Semantic Evaluation of Sign Language Translation

Model ReleasesDGX agent

arXiv:2606.03788v1 Announce Type: new Abstract: Sign Language Translation (SLT) is typically evaluated with surface-form metrics such as BLEU and ROUGE, which reward lexical overlap but do not directl

Social Caption: Evaluating Social Understanding in Multimodal Models

ResearchDGX agent

arXiv:2601.14569v2 Announce Type: replace Abstract: Social understanding abilities are crucial for multimodal large language models (MLLMs) to interpret human social interactions. We introduce SOCIAL

Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments

AgentsDGX agent

arXiv:2606.03892v1 Announce Type: cross Abstract: Training LLMs to orchestrate multi-step tool calls is held back by three coupled obstacles: realistic stateful execution environments are costly to bu

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2512.22539v2 Announce Type: replace-cross Abstract: While Vision-Language-Action models (VLAs) are rapidly advancing towards generalist robot policies, it remains difficult to quantitatively und

What’s new in serverless Managed Service for Apache Spark

HardwareDGX agent

Whether you use it for data preparation, real-time interactive queries, AI model training, or something entirely different, running Apache Spark at scale is demanding — you shouldn’t have to manage th

2 Jun 2026

3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

Model ReleasesDGX agent

arXiv:2606.01057v1 Announce Type: cross Abstract: Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neur

An Asynchronous Two-Speed Kalman Filter for Real-Time UUV Cooperative Navigation Under Acoustic Delays

AgentsDGX agent

arXiv:2604.02878v2 Announce Type: replace Abstract: In Global Navigation Satellite System (GNSS)-denied underwater environments, individual unmanned underwater vehicles (UUVs) suffer from unbounded de

An explainable hierarchical self attention-based approach for tremor detection in the time domain

TutorialsDGX agent

arXiv:2606.00461v1 Announce Type: new Abstract: Tremor is a common movement disorder associated with conditions like Parkinson's disease and Essential tremor, traditionally diagnosed through expert cl

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation

Model ReleasesDGX agent

arXiv:2606.00987v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temp

ASE-26: a curriculum for agentic software engineering as a discipline

Model ReleasesDGX agent

arXiv:2606.01152v1 Announce Type: cross Abstract: The work of a professional software engineer has begun to consist, increasingly, of directing agents rather than writing code, and the empirical evide

Attention mechanisms and transfer learning for robust peach leaf damage classification under domain shift

Model ReleasesDGX agent

arXiv:2606.02045v1 Announce Type: cross Abstract: Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in agricultural

AutoForest: Automatically Generating Forest Plots from Biomedical Studies with End-to-End Evidence Extraction and Synthesis

ApplicationsDGX agent

arXiv:2606.02403v1 Announce Type: cross Abstract: Systematic reviews rely on forest plots to synthesise quantitative evidence across biomedical studies, but generating them remains a fragmented and la

AutoIQ: An Ensemble Framework for Automatic Assessment of Geometric Distortion in Prostate Diffusion-Weighted Imaging

Local AiDGX agent

arXiv:2606.00393v1 Announce Type: cross Abstract: Geometric distortion in prostate diffusion-weighted imaging (DWI) can impair lesion localization and reduce the reliability of MRI-based clinical asse

Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation

SafetyDGX agent

arXiv:2602.11790v2 Announce Type: replace Abstract: Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in

BlueME: Robust Underwater Robot-to-Robot Communication Using Compact Magnetoelectric Antennas

AgentsDGX agent

arXiv:2411.09241v5 Announce Type: replace Abstract: We present the design, development, and experimental validation of BlueME, a compact magnetoelectric (ME) antenna array system for underwater robot-

Bridging Topology and Deep Representation Learning: A TDA-ViT Fusion Model for Four-Class Brain Tumor Classification

ResearchDGX agent

arXiv:2606.00927v1 Announce Type: new Abstract: Accurate brain tumor classification from magnetic resonance imaging (MRI) is a key requirement for early diagnosis and clinical decision-making. Vision

Can LLMs Reason Structurally? Benchmarking via the Lens of Data Structures

Model ReleasesDGX agent

arXiv:2505.24069v4 Announce Type: replace-cross Abstract: Large language models (LLMs) are deployed on increasingly complex tasks that require multi-step decision-making. Understanding their algorithm

Citation Grounding: Detecting and Reducing LLM Citation Hallucinations via Legal Citation Graphs

Model ReleasesDGX agent

arXiv:2606.00898v1 Announce Type: new Abstract: Large language models systematically hallucinate legal citations -- fabricating statute references, citing repealed provisions, and confusing jurisdicti

ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree

Model ReleasesDGX agent

arXiv:2606.01494v1 Announce Type: cross Abstract: Agent skills extend AI agents with reusable instructions, tools, scripts, references, and workflows, establishing a security boundary distinct from bo

Codex for every role, tool, and workflow

Model ReleasesDGX agent

OpenAI's Codex is an AI system that translates natural language instructions into code, designed to assist users across different roles, tools, and workflows. It enables developers, non-technical user

Codex is becoming a productivity tool for everyone

Model ReleasesDGX agent

OpenAI's Codex is evolving beyond code generation to become a general productivity tool accessible to non-programmers for knowledge work tasks. The tool leverages large language models to assist with

Constitutional Black-Box Monitoring for Scheming in LLM Agents

AgentsDGX agent

arXiv:2603.00829v2 Announce Type: replace-cross Abstract: Safe deployment of Large Language Model (LLM) agents in autonomous settings requires reliable oversight mechanisms. A central challenge is det

Context Matters: Repository-Aware Security Analysis of the Agent Skill Ecosystem

Model ReleasesDGX agent

arXiv:2603.16572v2 Announce Type: replace-cross Abstract: Agent skills extend local AI agents, such as Claude Code and OpenClaw, with additional functionality. Their growing popularity has led to dedi

CountGD++: Generalized Prompting for Open-World Counting

AgentsDGX agent

arXiv:2512.23351v2 Announce Type: replace Abstract: The flexibility and accuracy of methods for automatically counting objects in images and videos are limited by the way the object can be specified.

Cross-Generational Transfer of Adversarial Attacks Reveals Non-Monotonic Safety Alignment in LLMs

Model ReleasesDGX agent

arXiv:2606.00813v1 Announce Type: cross Abstract: Safety alignment in LLMs does not improve monotonically across model generations. Studying four generations of Google's Gemma family (7B-31B) with qua

DetailMaster: Can Your Text-to-Image Model Handle Long Prompts?

Model ReleasesDGX agent

arXiv:2505.16915v3 Announce Type: replace-cross Abstract: While recent Text-to-Image (T2I) models show impressive capabilities in synthesizing images from brief descriptions, they struggle with the lo

Early Prediction of Liver Cirrhosis Up to Two Years in Advance: A Machine Learning Study Benchmarking Against the FIB-4 and APRI Scores

Model ReleasesDGX agent

arXiv:2601.00175v2 Announce Type: replace Abstract: Objective: Develop and evaluate machine learning (ML) models for predicting incident liver cirrhosis (LC) one and two years prior to diagnosis using

Edge-Based QoS-Aware Adaptive Task Placement: A Closed-Loop Control in Multi-Robot Systems

ResearchDGX agent

arXiv:2606.00552v1 Announce Type: cross Abstract: Multi-robot systems (MRS) increasingly offload compute-intensive perception tasks to edge nodes to meet strict time-sensitive Quality-of-Service (QoS)

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning

Model ReleasesDGX agent

arXiv:2604.03893v2 Announce Type: replace Abstract: Current multimodal benchmarks for scientific reasoning primarily evaluate local information extraction -- models recognize symbols and values and th

FigSIM: A Dataset for Fine-grained Suicide Severity and Figurative Language in Suicide Memes

Model ReleasesDGX agent

arXiv:2606.02523v1 Announce Type: new Abstract: Suicide memes are memes used to express suicide-related thoughts or comment on suicide-related issues. Suicide memes are increasingly common on social m

FVSpec: Real-World Property-Based Tests as Lean Challenges

Model ReleasesDGX agent

arXiv:2606.01008v1 Announce Type: cross Abstract: We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first scrape 11,039 property-based tes

Hierarchical Online Prompt Mutation with Dual-Loop Feedback for Guardrailed Evidence Document Generation: A Production-Evaluation Case Study

Model ReleasesDGX agent

arXiv:2606.01472v1 Announce Type: cross Abstract: High-stakes production document-generation systems require language models to be adaptive, evidence-grounded, and auditable. We present HOPM, a hierar

HLL: Can Agents Cross Humanity's Last Line of Verification?

Model ReleasesDGX agent

arXiv:2606.02449v1 Announce Type: new Abstract: Multimodal agents are increasingly expected to operate interfaces on behalf of users, raising a central deployment question: can they truly substitute f

How Baz improved its AI Agent Code Review accuracy using Amazon Bedrock AgentCore

AgentsDGX agent

This post walks through how Baz built their Spec Review agent using Amazon Bedrock and Amazon Bedrock AgentCore. We'll cover the architecture decisions, implementation details, and the business outcom

Human in the Loop Adaptive Optimization for Improved Time Series Forecasting

TutorialsDGX agent

arXiv:2505.15354v2 Announce Type: replace Abstract: Time series forecasting models often produce systematic, predictable errors even in critical domains such as energy, finance, and healthcare. We int

Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning

SafetyDGX agent

arXiv:2606.00334v1 Announce Type: cross Abstract: Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

Model ReleasesDGX agent

arXiv:2606.01703v1 Announce Type: cross Abstract: We address the challenge of generating high-fidelity, long-form soundtracks that remain coherent across scene transitions. Existing AI music systems a

LLM agents patch security bugs, pass all tests, but still leave the vulnerability open [R]

ResearchDGX agent

Research demonstrates that LLM-based agents can generate functionally correct patches that pass all tests while still containing security vulnerabilities, challenging the assumption that test-passing

LLM Consortium for Software Design Refinement: A Controlled Experiment on Multi-Agent Collaboration Topologies

Model ReleasesDGX agent

arXiv:2606.01490v1 Announce Type: cross Abstract: We present a controlled experiment evaluating 12 multi-agent LLM collaboration topologies for software architecture design. Using a 2imes2imes2 factor

Monitoring Agentic Systems Before They're Reliable

AgentsDGX agent

arXiv:2606.02494v1 Announce Type: cross Abstract: Agentic systems entering production typically operate as partially integrated assemblies where structural defects, not task-level errors, dominate the

Not What, But How: A Communicative Audit of LLM Response Framing

ResearchDGX agent

arXiv:2606.02493v1 Announce Type: new Abstract: Large language models (LLMs) are being increasingly used to answer subjective, information-seeking questions, where users are sensitive to how responses

Order within Chaos: Capturing Intrinsic Energy Anomalies for AI-Manipulated Image Forgery Localization

Model ReleasesDGX agent

arXiv:2606.02178v1 Announce Type: cross Abstract: Recent advancements in generative AI have led to image editing models capable of producing realistic forgeries that evade traditional image forgery lo

Peacemaker at ATE-IT: Automatic term extraction from Italian text for waste management data using encoder model

ResearchDGX agent

arXiv:2606.01469v1 Announce Type: new Abstract: The development of automatic term extraction has become increasingly important in modern technology. Automatic term extraction can be found in virtually

Perspective on Bias in Biomedical AI: Preventing Downstream Healthcare Disparities

SafetyDGX agent

arXiv:2604.14514v2 Announce Type: replace Abstract: Healthcare disparities persist across socioeconomic boundaries, often attributed to unequal access to screening, diagnostics, and therapeutics. Howe

← Previous
1…6061626364…83
Next →