AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,910 results
Model Releases

Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

DGX agent

arXiv:2604.12371v1 Announce Type: new Abstract: We study typographic prompt injection attacks on vision-language models (VLMs), where adversarial text is rendered as images to bypass safety mechanisms

model-releasesarxiv-cs-cv
15 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss

DGX agent

arXiv:2604.12911v1 Announce Type: cross Abstract: Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to p

model-releasesarxiv-cs-ai
15 Apr 2026
Safety

Scaffold-Conditioned Preference Triplets for Controllable Molecular Optimization with Large Language Models

DGX agent

arXiv:2604.12350v1 Announce Type: cross Abstract: Molecular property optimization is central to drug discovery, yet many deep learning methods rely on black-box scoring and offer limited control over

safetyarxiv-cs-ai
15 Apr 2026
Research

Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling

DGX agent

arXiv:2508.04282v3 Announce Type: replace Abstract: Recent benchmarks for memory-augmented reinforcement learning (RL) have introduced partially observable Markov decision process (POMDP) environments

researcharxiv-cs-ai
15 Apr 2026
Safety

TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs

DGX agent

arXiv:2604.12232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed across diverse domains, yet their vulnerability to jailbreak attacks, where adversarial inputs

safetyarxiv-cs-ai
15 Apr 2026
Industry

This sub is mostly screenshots of ChatGPT being wrong. Meanwhile people on RunLobster (OpenClaw) are quietly running real businesses on AI.

DGX agent

This Reddit post contrasts the r/ChatGPT community's tendency to focus on AI failures and viral screenshots with a more pragmatic user base building real workflows on RunLobster, a managed cloud platf

industryr-chatgpt
15 Apr 2026
Industry

8 AI and data trends shaping financial services in 2026

DGX agent

Databricks outlines eight key AI and data trends expected to transform financial services in 2026, likely covering advancements in generative AI, real-time data processing, machine learning for risk m

industrydatabricks
14 Apr 2026
Safety

AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence

DGX agent

arXiv:2604.10579v1 Announce Type: cross Abstract: Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often constrained by geometric variations

safetyarxiv-cs-ai
14 Apr 2026
Industry

Big lab leaks

DGX agent

AI development is shifting toward fully integrated, agent-enabled applications, with major platforms like Anthropic's Claude and OpenAI

industryben-s-bites
14 Apr 2026
Safety

CoSToM:Causal-oriented Steering for Intrinsic Theory-of-Mind Alignment in Large Language Models

DGX agent

arXiv:2604.10031v1 Announce Type: cross Abstract: Theory of Mind (ToM), the ability to attribute mental states to others, is a hallmark of social intelligence. While large language models (LLMs) demon

safetyarxiv-cs-ai
14 Apr 2026
Safety

Do LLMs Know Tool Irrelevance? Demystifying Structural Alignment Bias in Tool Invocations

DGX agent

arXiv:2604.11322v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated impressive capabilities in utilizing external tools. In practice, however, LLMs are often exposed to to

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

ExecTune: Effective Steering of Black-Box LLMs with Guide Models

DGX agent

arXiv:2604.09741v1 Announce Type: cross Abstract: For large language models deployed through black-box APIs, recurring inference costs often exceed one-time training costs. This motivates composed age

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

FlashMem: Distilling Intrinsic Latent Memory via Computation Reuse

DGX agent

arXiv:2601.05505v2 Announce Type: replace Abstract: The stateless architecture of Large Language Models inherently lacks the mechanism to preserve dynamic context, compelling agents to redundantly rep

model-releasesarxiv-cs-cl
14 Apr 2026
Tools

Had a fantastic time at the @aiDotEngineer Europe conference in London last week. There was a lot of good coverage of people's emerging appr…

DGX agent

Had a fantastic time at the @aiDotEngineer Europe conference in London last week. There was a lot of good coverage of people's emerging approaches for working with AI coding agents. A couple of the st

toolsswyx--x
14 Apr 2026
Industry

How some mathematicians are exploring ways to incorporate LLM models into their research without losing direct experience with mathematical understanding (Konstantin Kakaes/Quanta Magazine)

DGX agent

Konstantin Kakaes / Quanta Magazine: How some mathematicians are exploring ways to incorporate LLM models into their research without losing direct experience with mathematical understanding — Those c

industrytechmeme
14 Apr 2026
Safety

However, most alignment research is not very crisp and requires research taste when evaluating. This is why we chose to point the AAR at thi…

DGX agent

However, most alignment research is not very crisp and requires research taste when evaluating. This is why we chose to point the AAR at this scalable oversight problem! Progress would let AARs work o

safetyjan-leike--x
14 Apr 2026
Safety

Learning to Focus: CSI-Free Hierarchical MARL for Reconfigurable Reflectors

DGX agent

arXiv:2604.05165v2 Announce Type: replace Abstract: Reconfigurable Intelligent Surfaces (RIS) has a potential to engineer smart radio environments for next-generation millimeter-wave (mmWave) networks

safetyarxiv-cs-ai
14 Apr 2026
Safety

Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit

DGX agent

arXiv:2604.09998v1 Announce Type: cross Abstract: Large language models (LLMs) have recently emerged as promising tools for augmenting Security Operations Center (SOC) workflows, with vendors increasi

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

Panoptic Pairwise Distortion Graph

DGX agent

arXiv:2604.11004v1 Announce Type: cross Abstract: In this work, we introduce a new perspective on comparative image assessment by representing an image pair as a structured composition of its regions.

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

ParseBench is the most comprehensive OCR benchmark for real-world enterprise documents: financial filings, contracts, insurance documents, a…

DGX agent

ParseBench is the most comprehensive OCR benchmark for real-world enterprise documents: financial filings, contracts, insurance documents, and more. We evaluate across 5 dimensions that are present am

model-releasesjerry-liu--x
14 Apr 2026
Hardware

Real-Time Voicemail Detection in Telephony Audio Using Temporal Speech Activity Features

DGX agent

arXiv:2604.09675v1 Announce Type: cross Abstract: Outbound AI calling systems must distinguish voicemail greetings from live human answers in real time to avoid wasted agent interactions and dropped c

hardwarearxiv-cs-ai
14 Apr 2026
Research

Retrieval Is Not Enough: Why Organizational AI Needs Epistemic Infrastructure

DGX agent

arXiv:2604.11759v1 Announce Type: new Abstract: Organizational knowledge used by AI agents typically lacks epistemic structure: retrieval systems surface semantically relevant content without distingu

researcharxiv-cs-ai
14 Apr 2026
Safety

Risk Awareness Injection: Calibrating Vision-Language Models for Safety without Compromising Utility

DGX agent

arXiv:2602.03402v3 Announce Type: replace Abstract: Vision language models (VLMs) extend the reasoning capabilities of large language models (LLMs) to cross-modal settings, yet remain highly vulnerabl

safetyarxiv-cs-ai
14 Apr 2026
Model Releases

Scaling unstructured enterprise knowledge with BigQuery Graph, and Kineviz GraphXR

DGX agent

Over 80% of enterprise data lives in unstructured form — PDFs, emails, reports, regulatory filings. Most of the time, such sources contain critical business information, yet they remain difficult to a

model-releasesgoogle-cloud-ai
14 Apr 2026
Model Releases

SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence

DGX agent

arXiv:2505.17012v3 Announce Type: replace-cross Abstract: Existing evaluations of multimodal large language models (MLLMs) on spatial intelligence are typically fragmented and limited in scope. In thi

model-releasesarxiv-cs-ai
14 Apr 2026
Tools

Start building a Sentry automation with a template from our marketplace: http://cursor.com/marketplace/automations/investigate-sentry-issues

DGX agent

Cursor offers a pre-built automation template in its marketplace that helps developers investigate Sentry issues more efficiently. The template, accessible at cursor.com/marketplace/automations/invest

toolscursor--x
14 Apr 2026
Model Releases

StarVLA-alpha: Reducing Complexity in Vision-Language-Action Systems

DGX agent

arXiv:2604.11757v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for building general-purpose robotic agents. However, the VLA landsc

model-releasesarxiv-cs-ai
14 Apr 2026
Safety

Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape

DGX agent

arXiv:2604.11184v1 Announce Type: cross Abstract: Context: Software engineering (SE) researchers increasingly study Generative AI (GenAI) while also incorporating it into their own research practices.

safetyarxiv-cs-ai
14 Apr 2026
Tutorials

The user interface of the future is your voice: Inside 8×8’s AI Studio

DGX agent

For years, the promise of “no-code” artificial intelligence has felt a bit like a “some assembly required” IKEA desk — sure, you aren’t sawing the wood yourself, but you’re trying to decipher how to a

tutorialssiliconangle
14 Apr 2026
Model Releases

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models

DGX agent

arXiv:2604.10733v1 Announce Type: cross Abstract: Large language models increasingly serve as conversational agents that adopt personas and role-play characters at user request. This capability, while

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean

DGX agent

arXiv:2604.08595v1 Announce Type: cross Abstract: Existing evaluation methods for LLM-based AI systems, such as LLM-as-a-Judge, verdict systems, and NLI, do not always align well with human assessment

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Benchmarked @DJLougen ’s Ornstein-27B-v2 Q6_K on my RTX 3090 using hermes-bench, my new open-source benchmarking UI for local LLMs and Herme…

DGX agent

Benchmarked @DJLougen ’s Ornstein-27B-v2 Q6_K on my RTX 3090 using hermes-bench, my new open-source benchmarking UI for local LLMs and Hermes agents. Ornstein is a Qwen 3.5 27B fine-tune trained on re

model-releasesclem-delangue--x
13 Apr 2026
Research

CaRLi-V: Camera-RADAR-LiDAR Point-Wise 3D Velocity Estimation

DGX agent

arXiv:2511.01383v2 Announce Type: replace Abstract: Accurate point-wise velocity estimation in 3D is crucial for robot interaction with non-rigid dynamic agents, enabling robust performance in path pl

researcharxiv-cs-ro
13 Apr 2026
Model Releases

China has erased the US lead in AI, Stanford HAI’s 2026 AI index reveals

DGX agent

Stanford University researchers today released their highly anticipated 2026 AI Index Report, revealing a global landscape where artificial intelligence technology is being adopted at record-breaking

model-releasessiliconangle
13 Apr 2026
Safety

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs

DGX agent

arXiv:2604.08846v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have been shown to be vulnerable to malicious queries that can elicit unsafe responses. Recent work uses prom

safetyarxiv-cs-ai
13 Apr 2026
Model Releases

From Business Events to Auditable Decisions: Ontology-Governed Graph Simulation for Enterprise AI

DGX agent

arXiv:2604.08603v1 Announce Type: new Abstract: Existing LLM-based agent systems share a common architectural failure: they answer from the unrestricted knowledge space without first simulating how ac

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

From Paper to Program: Accelerating Quantum Many-Body Algorithm Development via a Multi-Stage LLM-Assisted Workflow

DGX agent

arXiv:2604.04089v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can generate code rapidly but remain unreliable for scientific algorithms whose correctness depends on structural

model-releasesarxiv-cs-ai
13 Apr 2026
Safety

GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking

DGX agent

arXiv:2604.09222v1 Announce Type: cross Abstract: Audio large language models (ALLMs) enable rich speech-text interaction, but they also introduce jailbreak vulnerabilities in the audio modality. Exis

safetyarxiv-cs-ai
13 Apr 2026
Safety

I am catching glimpses in my feed that there is a backlash against Mythos as 'marketing hype,' and it is a little confusing. I don't think a…

DGX agent

I am catching glimpses in my feed that there is a backlash against Mythos as 'marketing hype,' and it is a little confusing. I don't think anyone who has used the latest agentic coding tools, would th

safetyethan-mollick--x
13 Apr 2026
Safety

Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism

DGX agent

arXiv:2604.09544v1 Announce Type: cross Abstract: Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely

safetyarxiv-cs-ai
13 Apr 2026
Safety

Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection

DGX agent

arXiv:2604.09024v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) have emerged as powerful tools for analyzing Internet-scale image data, offering significant benefits but al

safetyarxiv-cs-ai
13 Apr 2026
Safety

MAB-DQA: Addressing Query Aspect Importance in Document Question Answering with Multi-Armed Bandits

DGX agent

arXiv:2604.08952v1 Announce Type: new Abstract: Document Question Answering (DQA) involves generating answers from a document based on a user's query, representing a key task in document understanding

safetyarxiv-cs-cl
13 Apr 2026
Model Releases

Mitigating Extrinsic Gender Bias for Bangla Classification Tasks

DGX agent

arXiv:2411.10636v2 Announce Type: replace-cross Abstract: In this study, we investigate extrinsic gender bias in Bangla pretrained language models, a largely underexplored area in low-resource languag

model-releasesarxiv-cs-ai
13 Apr 2026
Safety

Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization

DGX agent

arXiv:2604.09253v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are powerful but remain vulnerable to multimodal jailbreak attacks. Existing attacks mainly rely on either explicit visu

safetyarxiv-cs-ai
13 Apr 2026
Model Releases

SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning

DGX agent

arXiv:2604.09452v1 Announce Type: cross Abstract: Safety guarantees are a prerequisite to the deployment of reinforcement learning (RL) agents in safety-critical tasks. Often, deployment environments

model-releasesarxiv-cs-ai
13 Apr 2026
Applications

SynDocDis: A Metadata-Driven Framework for Generating Synthetic Physician Discussions Using Large Language Models

DGX agent

arXiv:2604.08555v1 Announce Type: new Abstract: Physician-physician discussions of patient cases represent a rich source of clinical knowledge and reasoning that could feed AI agents to enrich and eve

applicationsarxiv-cs-cl
13 Apr 2026
Model Releases

The entire team put a lot of effort into this benchmark. Document parsing and OCR certainly isn't solved, but now it is a bit easier to meas…

DGX agent

The entire team put a lot of effort into this benchmark. Document parsing and OCR certainly isn't solved, but now it is a bit easier to measure 🦙 We’re open sourcing the first document OCR benchmark f

model-releasesjerry-liu--x
13 Apr 2026
Safety

Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation

DGX agent

arXiv:2604.09368v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly deployed as scalable user simulators for recommender system evaluation. Yet existing simulators per

safetyarxiv-cs-cv
13 Apr 2026
← Previous
1…300301302303304…374
Next →