AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,357 results
Model Releases

GAMBIT: A Gamified Jailbreak Framework for Multimodal Large Language Models

DGX agent

arXiv:2601.03416v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have become widely deployed, yet their safety alignment remains fragile under adversarial inputs. Previous

model-releasesarxiv-cs-cv
14 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Safety

Here's my first drive on Tesla FSD V14.3.1. It feels polished vs 14.3 and it my opinion is ready for wide release. • With 14.3.1, you can no…

DGX agent

Here's my first drive on Tesla FSD V14.3.1. It feels polished vs 14.3 and it my opinion is ready for wide release. • With 14.3.1, you can now tap on the new 'P' parking icon and bring up your differen

safetyelon-musk--x
14 Apr 2026
Safety

Hubble: An LLM-Driven Agentic Framework for Safe and Automated Alpha Factor Discovery

DGX agent

arXiv:2604.09601v1 Announce Type: new Abstract: Discovering predictive alpha factors in quantitative finance remains a formidable challenge due to the vast combinatorial search space and inherently lo

safetyarxiv-cs-ai
14 Apr 2026
Safety

Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs

DGX agent

arXiv:2604.10403v1 Announce Type: new Abstract: We address jailbreaks, backdoors, and unlearning for large language models (LLMs). Unlike prior work, which trains LLMs based on their actions when give

safetyarxiv-cs-lg
14 Apr 2026
Safety

LLM-as-Judge on a Budget

DGX agent

arXiv:2602.15481v2 Announce Type: replace Abstract: LLM-as-a-judge has emerged as a cornerstone technique for evaluating large language models by leveraging LLM reasoning to score prompt-response pair

safetyarxiv-cs-lg
14 Apr 2026
Safety

Minimal Embodiment Enables Efficient Learning of Number Concepts in Robot

DGX agent

arXiv:2604.11373v1 Announce Type: cross Abstract: Robots are increasingly entering human-interactive scenarios that require understanding of quantity. How intelligent systems acquire abstract numerica

safetyarxiv-cs-ai
14 Apr 2026
Safety

MM-LIMA: Less Is More for Alignment in Multi-Modal Datasets

DGX agent

arXiv:2308.12067v3 Announce Type: replace-cross Abstract: Multimodal large language models are typically trained in two stages: first pre-training on image-text pairs, and then fine-tuning using super

safetyarxiv-cs-ai
14 Apr 2026
Safety

MoRI: Mixture of RL and IL Experts for Long-Horizon Manipulation Tasks

DGX agent

arXiv:2604.10165v1 Announce Type: new Abstract: Reinforcement Learning (RL) and Imitation Learning (IL) are the standard frameworks for policy acquisition in manipulation. While IL offers efficient po

safetyarxiv-cs-ro
14 Apr 2026
Safety

🦔OpenAI is backing an Illinois state bill that would shield AI labs from liability in cases where their models cause mass casualties or lar…

DGX agent

🦔OpenAI is backing an Illinois state bill that would shield AI labs from liability in cases where their models cause mass casualties or large-scale financial disasters, defined as death or serious inj

safetygary-marcus--x
14 Apr 2026
Safety

Principles Do Not Apply Themselves: A Hermeneutic Perspective on AI Alignment

DGX agent

arXiv:2604.10673v1 Announce Type: new Abstract: AI alignment is often framed as the task of ensuring that an AI system follows a set of stated principles or human preferences, but general principles r

safetyarxiv-cs-ai
14 Apr 2026
Safety

Reasoning Resides in Layers: Restoring Temporal Reasoning in Video-Language Models with Layer-Selective Merging

DGX agent

arXiv:2604.11399v1 Announce Type: cross Abstract: Multimodal adaptation equips large language models (LLMs) with perceptual capabilities, but often weakens the reasoning ability inherited from languag

safetyarxiv-cs-cl
14 Apr 2026
Safety

Resilient Write: A Six-Layer Durable Write Surface for LLM Coding Agents

DGX agent

arXiv:2604.10842v1 Announce Type: cross Abstract: LLM-powered coding agents increasingly rely on tool-use protocols such as the Model Context Protocol~(MCP) to read and write files on a developer's wo

safetyarxiv-cs-ai
14 Apr 2026
Safety

Spatiotemporal-Aware Bit-Flip Injection on DNN-based Advanced Driver Assistance Systems (extended version)

DGX agent

arXiv:2604.03753v2 Announce Type: replace-cross Abstract: Modern advanced driver assistance systems (ADAS) rely on deep neural networks (DNNs) for perception and planning. Since DNNs' parameters resid

safetyarxiv-cs-lg
14 Apr 2026
Safety

Speaking to No One: Ontological Dissonance and the Double Bind of Conversational AI

DGX agent

arXiv:2604.10833v1 Announce Type: cross Abstract: Recent reports indicate that sustained interaction with conversational artificial intelligence (AI) systems can, in a small subset of users, contribut

safetyarxiv-cs-ai
14 Apr 2026
Safety

VLMaterial: Vision-Language Model-Based Camera-Radar Fusion for Physics-Grounded Material Identification

DGX agent

arXiv:2604.11671v1 Announce Type: cross Abstract: Accurate material recognition is a fundamental capability for intelligent perception systems to interact safely and effectively with the physical worl

safetyarxiv-cs-ro
14 Apr 2026
Safety

What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data

DGX agent

arXiv:2510.26202v2 Announce Type: replace-cross Abstract: Human feedback can alter language models in unpredictable and undesirable ways, as practitioners lack a clear understanding of what feedback d

safetyarxiv-cs-ai
14 Apr 2026
Safety

A new standard for research: How UC Riverside is securing the path to federal grants with Google Public Sector

DGX agent

At the University of California, Riverside (UCR), scientific breakthroughs depend on quickly moving from a hypothesis to a finished study. Yet for many researchers, the path to federal grants is often

safetygoogle-cloud-ai
13 Apr 2026
Safety

Anthropic says its $20M donation to Public First Action can't be 'used to influence federal elections' and is to educate the public on AI policy (Veronica Irwin/Transformer)

DGX agent

Veronica Irwin / Transformer: Anthropic says its $20M donation to Public First Action can't be “used to influence federal elections” and is to educate the public on AI policy — The company's money isn

safetytechmeme
13 Apr 2026
Safety

EGLOCE: Training-Free Energy-Guided Latent Optimization for Concept Erasure

DGX agent

arXiv:2604.09405v1 Announce Type: new Abstract: As text-to-image diffusion models grow increasingly prevalent, the ability to remove specific concepts-mostly explicit content and many copyrighted char

safetyarxiv-cs-cv
13 Apr 2026
Safety

@ESYudkowsky My horrendous nightmare of a political lifecycle, ladies and gentlemen and others.

DGX agent

Connor Leahy shared a post on X (formerly Twitter) quoting or referencing Eliezer Yudkowsky's account (@ESYudkowsky), describing what he characterizes as a 'horrendous nightmare of a political lifecyc

safetyconnor-leahy--x
13 Apr 2026
Safety

EthicMind: A Risk-Aware Framework for Ethical-Emotional Alignment in Multi-Turn Dialogue

DGX agent

arXiv:2604.09265v1 Announce Type: new Abstract: Intelligent dialogue systems are increasingly deployed in emotionally and ethically sensitive settings, where failures in either emotional attunement or

safetyarxiv-cs-cl
13 Apr 2026
Safety

Hierarchical Alignment: Enforcing Hierarchical Instruction-Following in LLMs through Logical Consistency

DGX agent

arXiv:2604.09075v1 Announce Type: new Abstract: Large language models increasingly operate under multiple instructions from heterogeneous sources with different authority levels, including system poli

safetyarxiv-cs-cl
13 Apr 2026
Safety

Import AI 453: Breaking AI agents; MirrorCode; and ten views on gradual disempowerment

DGX agent

Import AI issue 453 covers research and developments around vulnerabilities in AI agent systems, including methods for breaking or adversarially manipulating AI agents. The issue also features MirrorC

safetyimport-ai
13 Apr 2026
Safety

Learning Vision-Language-Action World Models for Autonomous Driving

DGX agent

arXiv:2604.09059v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have recently achieved notable progress in end-to-end autonomous driving by integrating perception, reasoning, and

safetyarxiv-cs-ai
13 Apr 2026
Safety

Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection

DGX agent

arXiv:2604.09024v1 Announce Type: cross Abstract: Multi-modal large language models (MLLMs) have emerged as powerful tools for analyzing Internet-scale image data, offering significant benefits but al

safetyarxiv-cs-ai
13 Apr 2026
Safety

LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving

DGX agent

arXiv:2604.08719v1 Announce Type: cross Abstract: Recent years have seen remarkable progress in autonomous driving, yet generalization to long-tail and open-world scenarios remains a major bottleneck

safetyarxiv-cs-ai
13 Apr 2026
Safety

Many Preferences, Few Policies: Towards Scalable Language Model Personalization

DGX agent

arXiv:2604.04144v2 Announce Type: replace-cross Abstract: The holy grail of LLM personalization is a single LLM for each user, perfectly aligned with that user's preferences. However, maintaining a se

safetyarxiv-cs-ai
13 Apr 2026
Safety

Mosaic: Multimodal Jailbreak against Closed-Source VLMs via Multi-View Ensemble Optimization

DGX agent

arXiv:2604.09253v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are powerful but remain vulnerable to multimodal jailbreak attacks. Existing attacks mainly rely on either explicit visu

safetyarxiv-cs-ai
13 Apr 2026
Safety

Scheming in the wild: detecting real-world AI scheming incidents with open-source intelligence

DGX agent

arXiv:2604.09104v1 Announce Type: cross Abstract: Scheming, the covert pursuit of misaligned goals by AI systems, represents a potentially catastrophic risk, yet scheming research suffers from signifi

safetyarxiv-cs-ai
13 Apr 2026
Safety

Verbalizing LLMs' assumptions to explain and control sycophancy

DGX agent

arXiv:2604.03058v2 Announce Type: replace-cross Abstract: LLMs can be socially sycophantic, affirming users when they ask questions like 'am I in the wrong?' rather than providing genuine assessment.

safetyarxiv-cs-ai
13 Apr 2026
Safety

⚡️ with @staysaasy, our second anonymous pod ever: https://www.youtube.com/watch?v=5KnCKadxSPY A conversation with the Stay Sassy duo on how…

DGX agent

⚡️ with @staysaasy, our second anonymous pod ever: https://www.youtube.com/watch?v=5KnCKadxSPY A conversation with the Stay Sassy duo on how AI is changing software teams, management, internal tooling

safetyswyx--x
13 Apr 2026
Safety

Each workflow in Thoth is a full LangGraph agent which can call subagents which themselves are LangGraph agents. Just describe what you want…

DGX agent

Each workflow in Thoth is a full LangGraph agent which can call subagents which themselves are LangGraph agents. Just describe what you want in plain English and it builds a full multi-step pipeline.

safetyharrison-chase--x
12 Apr 2026
Safety

🇧🇪Good news for Belgian Tesla owners! The Netherlands’ RDW just issued the first European type approval for Tesla FSD Supervised. Attached…

DGX agent

🇧🇪Good news for Belgian Tesla owners! The Netherlands’ RDW just issued the first European type approval for Tesla FSD Supervised. Attached letter from the Belgian federal administration (Minister Jean

safetyelon-musk--x
12 Apr 2026
Safety

Restore Britain would allow free public speech without imprisonment

DGX agent

Restore Britain would allow free public speech without imprisonment Platforms hosting lawful content must be shielded from government pressure to censor. We would require transparency in content moder

safetyelon-musk--x
12 Apr 2026
Safety

The Biden administration actively flew illegals into America with no vetting of their violent criminal past into America and paid for their …

DGX agent

The Biden administration actively flew illegals into America with no vetting of their violent criminal past into America and paid for their flights via NGOs. This is a war crime. Mayorkas and his budd

safetyelon-musk--x
12 Apr 2026
Safety

Under South Africa's Employment Equity policy, every employer with more than 50 staff must comply with RACIAL quota targets set by the gover…

DGX agent

Under South Africa's Employment Equity policy, every employer with more than 50 staff must comply with RACIAL quota targets set by the government. Under these rules, in roles such as 'skilled technici

safetyelon-musk--x
12 Apr 2026
Safety

Home safe!!!

DGX agent

I was unable to retrieve the specific content of the tweet at the URL provided (https://x.com/GaryMarcus/status/2042756801201082637). X (formerly Twitter) content is generally not accessible via we...

safetygary-marcus--x
11 Apr 2026
Safety

Homicides per 100,000 in El Salvador: 2015: 103 2016: 81.0 2017: 60.2 2018: 50.4 2019: 35.8 2020: 21.2 2021: 18.1 2022: 7.8 2023: 2.4 2024: …

DGX agent

El Salvador's homicide rate per 100,000 people declined dramatically from 103 in 2015 to 2.4 in 2023, representing a reduction of over 97% during that period. This sharp decline is widely attributed t

safetyelon-musk--x
11 Apr 2026
Safety

One investor today called for violence against me. Another lied about me, in a pretty deep and fundamental way. They are feeling the heat.

DGX agent

Gary Marcus, a prominent AI researcher and critic, posted on X (formerly Twitter) describing hostile reactions from investors, including one allegedly calling for violence against him and another maki

safetygary-marcus--x
11 Apr 2026
Safety

A Unified Multi-Layer Framework for Skill Acquisition from Imperfect Human Demonstrations

DGX agent

arXiv:2604.08341v1 Announce Type: new Abstract: Current Human-Robot Interaction (HRI) systems for skill teaching are fragmented, and existing approaches in the literature do not offer a cohesive frame

safetyarxiv-cs-ro
10 Apr 2026
Safety

Boycott OpenAI. They literally want the right to kill you.

DGX agent

I'm unable to fetch the content of that X (Twitter) URL directly, as I don't have the ability to browse or retrieve content from social media posts or URLs. Additionally, my web search did not retu...

safetygary-marcus--x
10 Apr 2026
Safety

Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judge

DGX agent

arXiv:2510.18196v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are commonly used as evaluators in various applications, but the reliability of the outcomes remains a challenge.

safetyarxiv-cs-ai
10 Apr 2026
Safety

CSA-Graphs: A Privacy-Preserving Structural Dataset for Child Sexual Abuse Research

DGX agent

arXiv:2604.07132v1 Announce Type: cross Abstract: Child Sexual Abuse Imagery (CSAI) classification is an important yet challenging problem for computer vision research due to the strict legal and ethi

safetyarxiv-cs-ai
10 Apr 2026
Safety

Data Leakage in Automotive Perception: Practitioners' Insights

DGX agent

arXiv:2604.06899v1 Announce Type: cross Abstract: Data leakage is the inadvertent transfer of information between training and evaluation datasets that poses a subtle, yet critical, risk to the reliab

safetyarxiv-cs-lg
10 Apr 2026
Safety

Deep Learning-Powered Visual SLAM Aimed at Assisting Visually Impaired Navigation

DGX agent

arXiv:2510.20549v2 Announce Type: replace Abstract: Despite advancements in SLAM technologies, robust operation under challenging conditions such as low-texture, motion-blur, or challenging lighting r

safetyarxiv-cs-cv
10 Apr 2026
Safety

Everything you need to know about “Open”AI’s claims to be working on AI “for the benefit of humanity”.

DGX agent

The specific X/Twitter post referenced (status ID 2042632799950352702) is not publicly accessible through search results, as X requires JavaScript and login to view individual posts. However, based...

safetygary-marcus--x
10 Apr 2026
Local Ai

FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization

DGX agent

arXiv:2604.06833v1 Announce Type: cross Abstract: As high quality public data becomes scarce, Federated Learning (FL) provides a vital pathway to leverage valuable private user data while preserving p

local-aiarxiv-cs-lg
10 Apr 2026
Safety

From experimentation to engagement: on the paradox of participatory AI and power in contexts of forced displacement and humanitarian crises

DGX agent

arXiv:2604.06219v1 Announce Type: cross Abstract: Across the Global North, calls for participatory artificial intelligence (AI) to improve the responsible, safe, and ethical use of AI have increased,

safetyarxiv-cs-ai
10 Apr 2026
← Previous
1…6263646566…300
Next →