AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “automated”

GridTimelineEvolution
4,983 results
5 Jun 2026

Seeking Counsel: Ongoing Targeted Campaign Against US Law Firms

SafetyDGX agent

Written by: Chad Reams, Tufail Ahmed, Keith Knapp, Ashley Frazer, Tyler McLellan Introduction From January through May 2026, Mandiant identified a financially motivated data theft extortion campaign e

4 Jun 2026

ALINC: Active Learning for Inductive Node Classification via Graph Sampling

Model ReleasesDGX agent

arXiv:2606.04647v1 Announce Type: new Abstract: Active learning (AL) for node classification typically focuses on selecting the most informative nodes for annotation within one or a few large graphs (

Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2407.03956v3 Announce Type: replace-cross Abstract: Prior research has enhanced the ability of Large Language Models (LLMs) to solve logic puzzles using techniques such as chain-of-thought promp

We have a Fleet agent called @docs_plz in our Slack that's made a very noticeable impact on our velocity of docs changes. In the chart below…

AgentsDGX agent

We have a Fleet agent called @docs_plz in our Slack that's made a very noticeable impact on our velocity of docs changes. In the chart below, you can clearly see that after it was added, the amount of

3 Jun 2026

Enhancing Paraphrase Type Generation: The Impact of DPO and RLHF Evaluated with Human-Ranked Data

ResearchDGX agent

arXiv:2506.02018v2 Announce Type: replace Abstract: Paraphrasing re-expresses meaning to enhance applications like text simplification, machine translation, and question-answering. Specific paraphrase

Lexicons and grammars for language processing: industrial or handcrafted products?

ResearchDGX agent

arXiv:2606.03412v1 Announce Type: new Abstract: During the recent years, the use of linguistic data for language processing increased progressively. Such data are now commonly called language resource

2 Jun 2026

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

Model ReleasesDGX agent

arXiv:2606.00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or

AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science

Model ReleasesDGX agent

arXiv:2603.19005v2 Announce Type: replace-cross Abstract: Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large

Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages

Model ReleasesDGX agent

arXiv:2606.00154v1 Announce Type: cross Abstract: Recent advancements in multimodal large language models (MLLMs) have achieved remarkable progress in multimodal reasoning and code generation, catalyz

Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs

Model ReleasesDGX agent

arXiv:2603.24511v2 Announce Type: replace-cross Abstract: We show that AI agents are capable of discovering novel algorithms for adversarial attacks against LLMs, advancing the state of the art on whi

CodeCytos: AI-assisted spatial molecular imaging analysis via code-augmented agent action space

Model ReleasesDGX agent

arXiv:2606.00472v1 Announce Type: cross Abstract: Conventional tissue image analysis software provides foundational capabilities for cellular analysis, including segmentation, basic morphological feat

Connecting AI agents with unstructured data using Google Cloud Storage MCP Servers

Model ReleasesDGX agent

Google Cloud Storage (GCS) is a foundational component of the modern agentic tech stack and the preferred home for unstructured data at scale. As enterprises deploy agents in production, the critical

RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates

SafetyDGX agent

arXiv:2506.11083v3 Announce Type: replace Abstract: We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate

RenoBench: A Citation Parsing Benchmark

Model ReleasesDGX agent

arXiv:2603.25640v2 Announce Type: replace-cross Abstract: Accurate parsing of citations is necessary for machine-readable scholarly infrastructure. But, despite sustained interest in this problem, exi

TalkTag: Fine-Grained Morphosyntactic Error Annotation for Transcribed Speech

ResearchDGX agent

arXiv:2606.01820v1 Announce Type: new Abstract: Fine-grained morphosyntactic error annotation is important in clinical and developmental language research, yet it is labour-intensive, expert-dependent

1 Jun 2026

AutoSci: A Memory-Centric Agentic System for the Full Scientific Research Lifecycle

AgentsDGX agent

arXiv:2605.31468v1 Announce Type: new Abstract: Scientific research has traditionally been human-intensive, requiring researchers to coordinate literature, ideas, experiments, manuscripts, and review

NEMO: Execution-Aware Optimization Modeling via Autonomous Coding Agents

AgentsDGX agent

arXiv:2601.21372v2 Announce Type: replace Abstract: We present NEMO, a system that translates Natural-language descriptions of decision problems into formal Executable Mathematical Optimization implem

Organizational Adaptation to Generative AI in Cybersecurity

SafetyDGX agent

arXiv:2506.12060v2 Announce Type: replace-cross Abstract: Cybersecurity organizations are adapting to GenAI integration through modified frameworks and hybrid operational processes, with success influ

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy

Model ReleasesDGX agent

arXiv:2602.22971v2 Announce Type: replace Abstract: As LLMs achieved breakthroughs in general reasoning, their proficiency in specialized scientific domains reveals pronounced gaps in existing benchma

29 May 2026

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

SafetyDGX agent

arXiv:2605.28994v1 Announce Type: new Abstract: AI tools to support real world decision making must be able to build simulation models that inform their recommendations and render them interpretable.

Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations

AgentsDGX agent

arXiv:2605.29786v1 Announce Type: new Abstract: Reproducibility is fundamental to the scientific method, yet remains a critical challenge in machine learning. Contributing factors include underspecifi

Decentralized LLM-Driven Coordination of Acoustic Robots for Contactless Object Manipulation

ApplicationsDGX agent

arXiv:2605.29378v1 Announce Type: new Abstract: Natural language interfaces can simplify interaction with multi-robot systems, especially when non-expert users need to issue high-level commands. Acous

MEMENTO: Leveraging Web as a Learning Signal for Low-Data Domains

TutorialsDGX agent

arXiv:2605.29795v1 Announce Type: new Abstract: Real-world tasks often lack large labeled datasets, motivating extensive work on learning in low-data regimes. However, existing approaches such as few-

Motion-guided sparse correction enables expert-quality point tracking across diverse microscopy regimes

ApplicationsDGX agent

arXiv:2605.29220v1 Announce Type: new Abstract: Tracking the dynamics of non-canonical biological systems in microscopy videos remains a persistent challenge. Both classical and learning-based tracker

28 May 2026

Announcing the newest cohort of the Google for Startups Accelerator: Middle East, North Africa & Turkey

Model ReleasesDGX agent

Google’s mission is to organize the world’s information and make it universally accessible. In high-growth, technically ambitious markets like the Middle East, North Africa, and Türkiye (MENA-T), we f

CFDTwin: An open-source GUI and Python toolkit for POD-NN surrogate modeling of ANSYS Fluent simulations

Model ReleasesDGX agent

arXiv:2605.27725v1 Announce Type: cross Abstract: High-fidelity computational fluid dynamics (CFD) is widely used for thermal-fluid design, but repeated CFD solves remain expensive for design optimiza

REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading

ApplicationsDGX agent

arXiv:2605.27402v1 Announce Type: cross Abstract: Open-ended grading is central to equitable and personalized education, yet manual grading remains time-consuming and costly, underscoring the need for

The Shape of Reasoning: Topological Analysis of Reasoning Traces in Large Language Models

ResearchDGX agent

arXiv:2510.20665v3 Announce Type: replace Abstract: Evaluating the quality of reasoning traces from large language models remains understudied, labor-intensive, and unreliable: current practice relies

27 May 2026

Beyond the Data Mesh Illusion: Designing Modern AI-augmented Lakehouses to Bridge the Gap Between Theory and Practice

SafetyDGX agent

arXiv:2605.27131v1 Announce Type: cross Abstract: Enterprise data platforms face an enduring tension between domain self-service and holistic governance. The data mesh paradigm proposed decentralized

Strategies for Guiding LLMs to Use Software Design Patterns: A Case of Singleton

Model ReleasesDGX agent

arXiv:2605.26898v1 Announce Type: cross Abstract: Large Language Models (LLMs) can generate functional source code from natural-language prompts, but often fail to consistently follow higher-level arc

26 May 2026

Artificial Effort

ResearchDGX agent

arXiv:2605.23920v1 Announce Type: cross Abstract: Real-effort tasks, in which participants perform cognitively costly activities whose outcomes depend on actual performance, are widely used in experim

Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models

Model ReleasesDGX agent

arXiv:2605.26100v1 Announce Type: cross Abstract: Code review is a critical practice in software engineering, yet the growing scale and frequency of code patches in modern projects, together with the

DeIDClinic: A Risk-Aware Pseudonymization Framework for Clinical Text De-identification and Re-identification Risk Assessment

ApplicationsDGX agent

arXiv:2410.01648v2 Announce Type: replace Abstract: The increasing availability of sensitive textual data has created an urgent need for robust de-identification methods that enable compliant data sha

Double Triangle Annotation: A Scalable Human-in-the-Loop Framework for High-Precision Historical Document Annotation

Model ReleasesDGX agent

arXiv:2605.25781v1 Announce Type: new Abstract: Evaluating structured-information extraction from historical documents at scale requires high-precision ground-truth annotations, yet traditional manual

Knowledge Graph Re-engineering Along the Ontological Continuum (extended version)

ApplicationsDGX agent

arXiv:2605.22093v2 Announce Type: replace Abstract: Knowledge graphs have become the primary vehicle for data integration and are critical to the success of modern AI, but the diversity of KG modellin

QUIET: A Multi-Blank Cascaded Story Cloze Benchmark for LLM Creative Generation Capability

Model ReleasesDGX agent

arXiv:2605.25955v1 Announce Type: cross Abstract: Large language models (LLMs) face a dual challenge in creative capability evaluation: existing benchmarks (e.g., Story Cloze Test, HellaSwag) measure

Reward Shaping and Action Masking for Compositional Tasks using Behavior Trees and LLMs

AgentsDGX agent

arXiv:2605.05795v2 Announce Type: replace Abstract: Decomposing complex tasks into a sequence of simpler subtasks can improve learning efficiency for an autonomous agent. Reinforcement learning (RL) c

RiskBridge: Turning CVEs into Business-Aligned Patch Priorities

SafetyDGX agent

arXiv:2601.06201v2 Announce Type: replace-cross Abstract: Enterprises are confronted with an unprecedented escalation in cybersecurity vulnerabilities, with thousands of new CVEs disclosed each month.

25 May 2026

A Proactive Multi-Agent Dialogue Framework for Assessing Social Language Disorder Traits in Autism

AgentsDGX agent

arXiv:2605.22993v1 Announce Type: cross Abstract: Characteristic linguistic behaviors associated with Social Language Disorder (SLD) in autism spectrum disorder, including echoic repetition, pronoun d

23 May 2026

Reinforced Graph of Thoughts: RL-Driven Adaptive Prompting for LLMs

ResearchDGX agent

arXiv:2605.22195v1 Announce Type: new Abstract: Graph of Thoughts (GoT), a generalized form of recent prompting paradigms for large language models (LLMs), has been shown to be useful for elaborate pr

21 May 2026

InternBootcamp Technical Report: Boosting LLM Reasoning with Verifiable Task Scaling

Model ReleasesDGX agent

arXiv:2508.08636v2 Announce Type: replace Abstract: Large language models (LLMs) have revolutionized artificial intelligence by enabling complex reasoning capabilities. While recent advancements in re

Towards Integrated Rock Support Visualisation in 3D Point Cloud of Underground Mines

ResearchDGX agent

arXiv:2605.20973v1 Announce Type: new Abstract: The effectiveness of rock support in underground mines depends on the interaction between installed rock bolts and the structural fabric of the surround

Validating Navmesh using Geometry: Voxel-Based Analysis with Prioritized Exploration

ApplicationsDGX agent

arXiv:2605.21397v1 Announce Type: cross Abstract: Navigation mesh (Navmesh) inconsistencies affect the player experience by directly impacting the navigation systems used by non-playable characters (N

20 May 2026

BLINKG: A Benchmark for LLM-Integrated Knowledge Graph Generation

Model ReleasesDGX agent

arXiv:2605.19518v1 Announce Type: new Abstract: Generating Knowledge Graphs (KGs) remains one of the most time-consuming and labor-intensive tasks for knowledge engineers, as they need to identify sem

MapAnything: Evaluating Monocular Metric Depth Models for 3D Urban Asset Localization

ResearchDGX agent

arXiv:2509.14839v2 Announce Type: replace Abstract: City administrations increasingly rely on comprehensive databases and urban digital twins of city assets, such as traffic signs and trees, as well a

Operationalising Artificial Intelligence Bills of Materials (AIBOMs) for Verifiable AI Provenance and Lifecycle Assurance

AgentsDGX agent

arXiv:2605.19755v1 Announce Type: cross Abstract: Artificial Intelligence (AI) systems are increasingly dependent on complex, multi-layered software supply chains that introduce challenges for reprodu

19 May 2026

ALIGN: A Vision-Language Framework for High-Accuracy Accident Location Inference through Geo-Spatial Neural Reasoning

Local AiDGX agent

arXiv:2511.06316v3 Announce Type: replace Abstract: In low- and middle-income countries, public safety and urban planning initiatives frequently face a critical shortage of accurate, location-specific

EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL

AgentsDGX agent

arXiv:2605.18703v1 Announce Type: new Abstract: Equipping LLMs with tool-use capabilities via Agentic Reinforcement Learning (Agentic RL) is bottlenecked by two challenges: the lack of scalable, robus

Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains

ResearchDGX agent

arXiv:2605.18261v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has demonstrated promising potential to enhance the reasoning capabilities of large language model

Multilingual jailbreaking of LLMs using low-resource languages

Model ReleasesDGX agent

arXiv:2605.18239v1 Announce Type: cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attempts that circumvent safety guardrails. We investigate whether multi-turn conversation

RL4RLA: Teaching ML to Discover Randomized Linear Algebra Algorithms Through Curriculum Design and Graph-Based Search

SafetyDGX agent

arXiv:2605.18004v1 Announce Type: new Abstract: Randomized linear algebra (RLA) algorithms are a modern class of numerical linear algebra techniques that play an essential role in scientific computing

Velocity and stroke rate reconstruction of canoe sprint team boats based on panned and zoomed video recordings

ResearchDGX agent

arXiv:2602.22941v2 Announce Type: replace Abstract: Pacing strategies, defined by velocity and stroke rate profiles, are essential for peak performance in canoe sprint. While GPS is the gold standard

18 May 2026

RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision

Model ReleasesDGX agent

arXiv:2605.15537v1 Announce Type: new Abstract: This paper introduces RTL-BenchMT, an agentic framework for dynamically maintaining RTL generation benchmarks. Large Language Models (LLMs) assisted aut

15 May 2026

Databricks brings GPT-5.5 to enterprise agent workflows

Model ReleasesDGX agent

Databricks has integrated OpenAI's GPT-5.5 model into enterprise agent workflows, enabling organizations to build and deploy AI agents with advanced language capabilities. This partnership leverages D

I strongly believe there are entire companies right now under heavy AI psychosis and its impossible to have rational conversations about it …

TutorialsDGX agent

I strongly believe there are entire companies right now under heavy AI psychosis and its impossible to have rational conversations about it with them. I can't name any specific people because they inc

Lang2MLIP: End-to-End Language-to-Machine Learning Interatomic Potential Development with Autonomous Agentic Workflows

AgentsDGX agent

arXiv:2605.14527v1 Announce Type: new Abstract: Developing machine learning interatomic potentials (MLIPs) for complex materials systems remains challenging because it requires expertise in atomistic

OPT-Engine: Benchmarking the Limits of LLMs in Optimization Modeling via Complexity Scaling

Model ReleasesDGX agent

arXiv:2601.19924v2 Announce Type: replace-cross Abstract: We investigate the capabilities and scalability of Large Language Models (LLMs) in optimization modeling, a domain requiring structured reason

14 May 2026

scShapeBench: Discovering geometry from high dimensional scRNAseq data

Model ReleasesDGX agent

arXiv:2605.12662v1 Announce Type: new Abstract: High-dimensional point cloud data arise across many scientific domains, especially single-cell biology. The shapes or topologies of these datasets deter

13 May 2026

FLAME: A New Dataset on FLemish Accounts of Momentary Experiences

Model ReleasesDGX agent

arXiv:2504.14707v3 Announce Type: replace Abstract: We introduce FLAME (FLemish Accounts of Momentary Experiences), a new corpus of nearly 25,000 daily personal narratives in Belgian-Dutch (Flemish),

The new era of SaMD: Why cloud infrastructure is the foundation for digital health in 2026

SafetyDGX agent

In the healthcare and life sciences industries, speed saves lives, but meeting regulatory requirements and other administrative burdens often pumps the brakes for manufacturers of software as a medica

← Previous
1…3839404142…84
Next →