AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,357 results
11 Apr 2026

Home safe!!!

SafetyDGX agent

I was unable to retrieve the specific content of the tweet at the URL provided (https://x.com/GaryMarcus/status/2042756801201082637). X (formerly Twitter) content is generally not accessible via we...

Homicides per 100,000 in El Salvador: 2015: 103 2016: 81.0 2017: 60.2 2018: 50.4 2019: 35.8 2020: 21.2 2021: 18.1 2022: 7.8 2023: 2.4 2024: …

SafetyDGX agent

El Salvador's homicide rate per 100,000 people declined dramatically from 103 in 2015 to 2.4 in 2023, representing a reduction of over 97% during that period. This sharp decline is widely attributed t

One investor today called for violence against me. Another lied about me, in a pretty deep and fundamental way. They are feeling the heat.

SafetyDGX agent

Gary Marcus, a prominent AI researcher and critic, posted on X (formerly Twitter) describing hostile reactions from investors, including one allegedly calling for violence against him and another maki

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
10 Apr 2026

A Unified Multi-Layer Framework for Skill Acquisition from Imperfect Human Demonstrations

SafetyDGX agent

arXiv:2604.08341v1 Announce Type: new Abstract: Current Human-Robot Interaction (HRI) systems for skill teaching are fragmented, and existing approaches in the literature do not offer a cohesive frame

Boycott OpenAI. They literally want the right to kill you.

SafetyDGX agent

I'm unable to fetch the content of that X (Twitter) URL directly, as I don't have the ability to browse or retrieve content from social media posts or URLs. Additionally, my web search did not retu...

Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge

SafetyDGX agent

arXiv:2510.18196v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are commonly used as evaluators in various applications, but the reliability of the outcomes remains a challenge.

CSA-Graphs: A Privacy-Preserving Structural Dataset for Child Sexual Abuse Research

SafetyDGX agent

arXiv:2604.07132v1 Announce Type: cross Abstract: Child Sexual Abuse Imagery (CSAI) classification is an important yet challenging problem for computer vision research due to the strict legal and ethi

Data Leakage in Automotive Perception: Practitioners' Insights

SafetyDGX agent

arXiv:2604.06899v1 Announce Type: cross Abstract: Data leakage is the inadvertent transfer of information between training and evaluation datasets that poses a subtle, yet critical, risk to the reliab

Deep Learning-Powered Visual SLAM Aimed at Assisting Visually Impaired Navigation

SafetyDGX agent

arXiv:2510.20549v2 Announce Type: replace Abstract: Despite advancements in SLAM technologies, robust operation under challenging conditions such as low-texture, motion-blur, or challenging lighting r

Everything you need to know about “Open”AI’s claims to be working on AI “for the benefit of humanity”.

SafetyDGX agent

The specific X/Twitter post referenced (status ID 2042632799950352702) is not publicly accessible through search results, as X requires JavaScript and login to view individual posts. However, based...

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization

Local AiDGX agent

arXiv:2604.06833v1 Announce Type: cross Abstract: As high quality public data becomes scarce, Federated Learning (FL) provides a vital pathway to leverage valuable private user data while preserving p

From experimentation to engagement: on the paradox of participatory AI and power in contexts of forced displacement and humanitarian crises

SafetyDGX agent

arXiv:2604.06219v1 Announce Type: cross Abstract: Across the Global North, calls for participatory artificial intelligence (AI) to improve the responsible, safe, and ethical use of AI have increased,

Front-End Ethics for Sensor-Fused Health Conversational Agents: An Ethical Design Space for Biometrics

SafetyDGX agent

arXiv:2604.06203v1 Announce Type: cross Abstract: The integration of continuous data from built-in sensors and Large Language Models (LLMs) has fueled a surge of 'Sensor-Fused LLM agents' for personal

Governing frontier general-purpose AI in the public sector: adaptive risk management and policy capacity under uncertainty through 2030

SafetyDGX agent

arXiv:2604.06215v1 Announce Type: cross Abstract: The governance of frontier general-purpose artificial intelligence has become a public-sector problem of institutional design, not merely a technical

Harnessing Embodied Agents: Runtime Governance for Policy-Constrained Execution

SafetyDGX agent

arXiv:2604.07833v1 Announce Type: new Abstract: Embodied agents are evolving from passive reasoning systems into active executors that interact with tools, robots, and physical environments. Once gran

Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization

SafetyDGX agent

arXiv:2604.06285v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have become essential for tasks such as image synthesis, captioning, and retrieval by aligning textual and visual inform

Incorporating Social Awareness into Control of Unknown Multi-Agent Systems: A Real-Time Spatiotemporal Tubes Approach

SafetyDGX agent

arXiv:2510.25597v2 Announce Type: replace-cross Abstract: This paper presents a decentralized control framework that incorporates social awareness into multi-agent systems with unknown dynamics to ach

Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents

SafetyDGX agent

arXiv:2510.07809v4 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) empower autonomous mobile agents, yet their security under realistic mobile deployment constraints remain

Learning Who Disagrees: Demographic Importance Weighting for Modeling Annotator Distributions with DiADEM

SafetyDGX agent

arXiv:2604.08425v1 Announce Type: cross Abstract: When humans label subjective content, they disagree, and that disagreement is not noise. It reflects genuine differences in perspective shaped by anno

LINE: LLM-based Iterative Neuron Explanations for Vision Models

SafetyDGX agent

arXiv:2604.08039v1 Announce Type: new Abstract: Interpreting the concepts encoded by individual neurons in deep neural networks is a crucial step towards understanding their complex decision-making pr

LUMINA: Foundation Models for Topology Transferable ACOPF

SafetyDGX agent

arXiv:2603.04300v2 Announce Type: replace Abstract: Foundation models in general promise to accelerate scientific computation by learning reusable representations across problem instances, yet constra

Mark “Metaverse” Zuckerberg totally bought Moltbook at the peak of the market 🤣

SafetyDGX agent

Meta acquired Moltbook, an AI-only social network launched in January 2026 by entrepreneurs Matt Schlicht and Ben Parr, after the platform had gone viral and then faded in popularity. Moltbook is ...

MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2604.08203v1 Announce Type: new Abstract: Medical Vision-Language Models (VLMs) hold immense promise for complex clinical tasks, but their reasoning capabilities are often constrained by text-on

Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability

SafetyDGX agent

arXiv:2604.06628v1 Announce Type: new Abstract: A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes. We revisit t

Safe Large-Scale Robust Nonlinear MPC in Milliseconds via Reachability-Constrained System Level Synthesis on the GPU

SafetyDGX agent

arXiv:2604.07644v1 Announce Type: new Abstract: We present GPU-SLS, a GPU-parallelized framework for safe, robust nonlinear model predictive control (MPC) that scales to high-dimensional uncertain rob

Towards the Development of an LLM-Based Methodology for Automated Security Profiling in Compliance with Ukrainian Cybersecurity Regulations

SafetyDGX agent

arXiv:2604.06274v1 Announce Type: cross Abstract: In recent years, the pace of development of information technology in various areas has increased drastically, forcing cybersecurity specialists to co

TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories

Model ReleasesDGX agent

arXiv:2604.07223v1 Announce Type: cross Abstract: As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final outputs to int

WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks

SafetyDGX agent

arXiv:2604.06367v1 Announce Type: cross Abstract: Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries. While current benchmarks evaluate

9 Apr 2026

All In’s @davidsacks liking one of my tweets was not in my 2026 bingo card.

SafetyDGX agent

The specific tweet from Gary Marcus (X post ID 2042361570253217968) is not publicly accessible without authentication, and the search results do not surface its specific content. Based on available...

GenAI’s popularity has hit a wall. Hard to see how OpenAI is going to make its numbers, and easy to see why they bought a media company. The…

SafetyDGX agent

AI critic Gary Marcus argues that GenAI's popularity has plateaued, with a landmark MIT study finding that only 5% of enterprise AI pilots generate revenue, while most deliver little to no measura...

Google's AI Overviews spew out millions of false answers per hour, bombshell study reveals https://trib.al/1ao7qB1

SafetyDGX agent

A study commissioned by *The New York Times* and conducted by AI startup Oumi tested 4,326 Google searches using the SimpleQA benchmark, finding that Google's AI Overviews were accurate 85% of the...

I will never get used to the sheer number of people who lie about me.

SafetyDGX agent

I was unable to retrieve results for that specific tweet URL. The tweet ID referenced (2042037797503299600) does not appear in any indexed web search results, and X (formerly Twitter) posts are gen...

Stargate never made sense.

SafetyDGX agent

AI critic and cognitive scientist Gary Marcus publicly challenged the economic rationale behind Project Stargate, the Trump-announced $500 billion AI infrastructure initiative, arguing that the mat...

this map is the story of the last 25 years of us foreign policy in its own backyard. and that was before the tariffs.

SafetyDGX agent

I was unable to retrieve the specific content of the linked X (Twitter) post by Ian Bremmer, as the URL points to a social media post that is not directly accessible or indexed with its full conten...

This SHOULD be an obvious point.

SafetyDGX agent

I was unable to retrieve the specific tweet at the URL provided (https://x.com/GaryMarcus/status/2042242813384142965). The tweet ID (2042242813384142965) appears to be from the future relative to c...

Tragically I am continuing to find that the most effective guardrail against slop is extremely talented engineers doing very thoughtful, hum…

SafetyDGX agent

I was unable to retrieve the specific content from the X (Twitter) post at the URL provided, and my web search did not surface the original post or any reliable secondary sources quoting or summari...

Trump's emergency orders pushing coal power are 'illegal' as well as dumb

SafetyDGX agent

The Trump administration's Department of Energy has invoked Section 202(c) of the Federal Power Act — a provision granting broad emergency authority over the electricity system that had previously...

8 Apr 2026

Again, if you care about computer security, read the red team report: https://red.anthropic.com/2026/mythos-preview/

SafetyDGX agent

Anthropic's Frontier Red Team report (April 2026) details the cybersecurity capabilities of Claude Mythos Preview, a general-purpose frontier model that performs strongly across the board but is s...

Common Failure Modes Break VLM-Powered OCR in Production. 🔁 Repetition Loops — model spirals into infinite whitespace, exhausts resources, …

SafetyDGX agent

Common Failure Modes Break VLM-Powered OCR in Production. 🔁 Repetition Loops — model spirals into infinite whitespace, exhausts resources, cascades latency across your system 🛑 Recitation Errors — saf

Dudes who won’t tag me because they know their arguments are weak sauce.* *for intellectual exercise you can list the flaws and misrepresent…

SafetyDGX agent

I was unable to retrieve the specific tweet at that URL — the X (Twitter) page requires JavaScript/login to load, and the tweet ID `2041954164562145434` does not appear in any indexed search result...

In AI, a lot can change in seven years.

SafetyDGX agent

The specific tweet (status ID 2041953155651661977) is not accessible — that ID appears to be from a future date and does not correspond to any retrievable post in the search results. The URL provid...

Incredible.

SafetyDGX agent

I was unable to retrieve the specific post at the URL provided (status ID `2041955770644713823`). This post ID does not appear in any search results, and X (formerly Twitter) requires JavaScript/lo...

literally fourteen minutes after my last explanation of why this is a false dichotomy 🤦‍♂️

SafetyDGX agent

The specific tweet (status ID 2041904683338625283) is not publicly accessible through search results, and the URL provided appears to reference a future or inaccessible post. The tweet ID is also b...

not surprised by any of this, headline or subheading

SafetyDGX agent

I was unable to retrieve the specific tweet at that URL (tweet ID 2041912293475414378). The tweet ID is extremely high — well beyond current Twitter/X ID ranges as of today — suggesting it may be a...

Voice ChatGPT can’t start a timer, but AGI is imminent! 🤦‍♂️

SafetyDGX agent

AI critic Gary Marcus uses the irony of ChatGPT's Voice mode being unable to perform a basic task — starting a timer — as a pointed illustration of the gap between AI industry hype and real-world c...

7 Apr 2026

You should read the red team report: https://red.anthropic.com/2026/mythos-preview/

SafetyDGX agent

Anthropic's Frontier Red Team published a technical report (April 2026) detailing how their unreleased model, Claude Mythos Preview, autonomously identifies and exploits critical security vulnerabi...

13 Aug 2026

A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation

SafetyDGX agent

arXiv:2512.06547v4 Announce Type: replace-cross Abstract: Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL setti

A method for tissue-mask supported whole-body image registration in the UK Biobank

SafetyDGX agent

arXiv:2512.02702v3 Announce Type: replace Abstract: The UK Biobank is a large-scale study collecting imaging and non-imaging health data. Robust and accurate inter-subject image registration of these

A New First-Order Meta-Learning Algorithm with Convergence Guarantees

SafetyDGX agent

arXiv:2409.03682v2 Announce Type: replace Abstract: Learning new tasks by leveraging prior experience is a fundamental trait of intelligent systems. While Model-Agnostic Meta-Learning (MAML) is a lead

Accuracy and Order Sensitivity Diverge Under Label-Free Strategies

SafetyDGX agent

arXiv:2608.11947v1 Announce Type: cross Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, whic

Adaptation of Generalist Robot Policies with Minimal Data

SafetyDGX agent

arXiv:2608.11363v1 Announce Type: cross Abstract: A central goal in robot learning is to move beyond task-specific human data collection toward robots that improve through autonomous interaction. Yet

Alignment of Similarity-Transformed Images Based on Fourier--Mellin Transform Using Auxiliary Function Method

SafetyDGX agent

arXiv:2608.11565v1 Announce Type: cross Abstract: This paper proposes an algorithm for estimating the similarity transformation, namely translation, scale, and rotation, between two images with subpix

anthropic’s best hope is that OpenAI falls apart. which it might.

SafetyDGX agent

anthropic’s best hope is that OpenAI falls apart. which it might. NEW from Ramp AI Index: disappointing adoption of Fable 5 We've heard several reasons from businesses...mainly Fable 5 is just too exp

Anti-Shortcut Distillation via Temporal Negative Knowledge Transfer

SafetyDGX agent

arXiv:2608.11789v1 Announce Type: new Abstract: Knowledge distillation (KD) trains a compact student by attracting it towards a converged teacher. It is silent about which directions the teacher itsel

AutoGrable: What Is a Good Graph for a Table?

SafetyDGX agent

arXiv:2608.11431v1 Announce Type: new Abstract: Graph learning presupposes a graph, and tables and relational databases do not come with one. Applying a GNN to them requires deciding which entities be

Benchmarking Trustworthiness of SLMs: Pre-trained vs. Compressed

SafetyDGX agent

arXiv:2608.11981v1 Announce Type: new Abstract: Small Language Models (SLMs) have emerged as a more efficient alternative to traditional Large Language Models (LLMs), offering promising potential in r

Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models

SafetyDGX agent

arXiv:2608.12078v1 Announce Type: cross Abstract: Learning world models from offline trajectories enables agents to accomplish different tasks through planning. Object-centric (OC) representations, wh

CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement

SafetyDGX agent

arXiv:2608.11631v1 Announce Type: new Abstract: In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In

CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications

SafetyDGX agent

arXiv:2608.11588v1 Announce Type: new Abstract: Mobile GUI agents remain brittle when deployed to applications absent from source training. We study novel-app generalization under a limited target int

Continuous-Latent Predictive Modeling with Semantic Alignment for EEG-Language Foundation Models

SafetyDGX agent

arXiv:2608.11656v1 Announce Type: new Abstract: Recent advances in EEG foundation models have demonstrated the potential of large-scale pretraining to enable generalizable neural decoding across subje

← Previous
1…5051525354…240
Next →