AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,812 results
8 Jun 2026

RAVEN: Retrieval-Augmented Vulnerability Exploration Network for Memory Corruption Analysis in User Code and Binary Programs

SafetyDGX agent

arXiv:2604.17948v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across various cybersecurity tasks, including vulnerability classificat

Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability

SafetyDGX agent

arXiv:2606.07437v1 Announce Type: cross Abstract: The ISO 26262 standard defines functional safety for road vehicles through risk assessments based on Severity, Exposure, and Controllability, grounded

Residual-Controlled Multiplier Learning for Stochastic Constrained Decision-Making

SafetyDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.07088v1 Announce Type: new Abstract: Stochastic constrained decision-making requires optimizing performance objectives while enforcing statistical requirements such as safety or fairness. H

Robotic Policy Adaptation via Weight-Space Meta-Learning

SafetyDGX agent

arXiv:2606.07217v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are emerging as a promising paradigm for robotic manipulation, enabling general-purpose policies trained from larg

Robots Need More than VLA and World Models

SafetyDGX agent

arXiv:2606.06556v1 Announce Type: new Abstract: Generalist robot intelligence is often framed as a policy-scaling problem: collect more robot demonstrations, train larger Vision-Language-Action (VLA)

Robust Driving Control for Autonomous Vehicles: An Intelligent General-sum Constrained Adversarial Reinforcement Learning Approach

SafetyDGX agent

arXiv:2510.09041v3 Announce Type: replace-cross Abstract: Deep reinforcement learning (DRL) has demonstrated remarkable success in developing autonomous driving policies. However, its vulnerability to

SafeGene: Reusable Adapters for Transferable Safety Alignment

SafetyDGX agent

arXiv:2606.06519v1 Announce Type: new Abstract: Open-weight LLMs are increasingly fine-tuned into customized assistants, but downstream fine-tuning can weaken safety alignment and make models more vul

Self-evolving LLM agents with in-distribution Optimization

SafetyDGX agent

arXiv:2606.07367v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently emerged as powerful controllers for interactive agents in complex environments, yet training them to perform

Semantic-Structural Alignment for Generative Pictorial Charts

SafetyDGX agent

arXiv:2606.06498v1 Announce Type: cross Abstract: Traditional statistical graphics are precise but often lack the visual appeal, memorability, and engagement of pictorial charts. We present a generati

SERNF: Sample-Efficient Real-World Dexterous Policy Fine-Tuning via Action-Chunked Critics and Normalizing Flows

SafetyDGX agent

arXiv:2602.09580v4 Announce Type: replace-cross Abstract: Real-world fine-tuning of dexterous manipulation policies remains challenging due to limited real-world interaction budgets and highly multimo

Shield-Loco: Shielding Locomotion Policies with Predictive Safety Filtering

SafetyDGX agent

arXiv:2606.07193v1 Announce Type: new Abstract: Reinforcement learning (RL) policies enable dynamic legged locomotion but lack mechanisms to avoid violations of safety constraints that are absent duri

Silverfort brings runtime identity controls to Microsoft Copilot Studio agents

SafetyDGX agent

Identity security company Silverfort Inc. today launched an integration that applies its identity and access controls to artificial intelligence agents built into Microsoft Corp.’s Copilot Studio, enf

Simulation-Driven Imitation Learning for Biosignals-Free Shared-Autonomy Prosthetic Grasping

SafetyDGX agent

arXiv:2606.07389v1 Announce Type: new Abstract: Biosignals-free shared-autonomy control of upper-limb prosthetic hands aims to enable natural and low-effort manipulation without relying on EMG or othe

SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating

SafetyDGX agent

arXiv:2606.07074v1 Announce Type: cross Abstract: Deep research agents have demonstrated remarkable capabilities in complex information-seeking tasks, yet this power comes at a steep computational cos

Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills

SafetyDGX agent

arXiv:2606.07412v1 Announce Type: cross Abstract: LLM-driven software engineering agents have become a central testbed for real-world language-model capability, yet their training remains limited by t

SpaceX goes public Friday at ~94x revenue. Across 45 years of data, IPOs that debut above 40x sales underperform the market by 58% over the …

SafetyDGX agent

SpaceX goes public Friday at ~94x revenue. Across 45 years of data, IPOs that debut above 40x sales underperform the market by 58% over the next 3 years, and by 76% style-adjusted. The golden rule of

Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry

SafetyDGX agent

arXiv:2603.26846v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) expand in capability and application scope, their trustworthiness becomes critical. A vital risk is intrinsic

Step-Wise Refusal Dynamics in Autoregressive and Diffusion Language Models

SafetyDGX agent

arXiv:2602.02600v3 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) have recently emerged as a competitive alternative to autoregressive (AR) models, offering parallel decoding,

SV-Detect: AI-generated Text Detection with Steering Vectors

SafetyDGX agent

arXiv:2606.07313v1 Announce Type: cross Abstract: Detecting machine-generated text is especially difficult under distribution shift, such as transfer across domains, source models, and editing attacks

Sycophantic Praise: Evaluating Excessive Praise in Language Models

SafetyDGX agent

arXiv:2606.07441v1 Announce Type: new Abstract: Sycophancy in language models is typically studied as excessive agreement or validation, while explicit praise and flattery have received comparatively

T-GMP: Terrain-conditioned Generative Motion Priors for Versatile and Natural Humanoid Locomotion

SafetyDGX agent

arXiv:2606.06944v1 Announce Type: new Abstract: Achieving both anthropomorphic naturalness and robust terrain traversal remains a fundamental challenge in humanoid locomotion. Existing Reinforcement L

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling

SafetyDGX agent

arXiv:2507.06419v3 Announce Type: replace Abstract: Reward modeling (RM), which captures human preferences to align large language models (LLMs), is increasingly employed in tasks such as model finetu

Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization

SafetyDGX agent

arXiv:2606.07000v1 Announce Type: new Abstract: Recent post-training methods, particularly Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced the reasoning ability of L

The discovery of the effects of women employment participation on the fertility of developing countries: A panel data approach

SafetyDGX agent

arXiv:2606.07093v1 Announce Type: new Abstract: The fertility trend in developing countries has experienced a significant decline in the last few decades; at the same time, the role of women in the wo

The legendary investor Vinod Khosla is in the news because of questions about his integrity. I want to share my own experience: he lied vici…

SafetyDGX agent

The legendary investor Vinod Khosla is in the news because of questions about his integrity. I want to share my own experience: he lied viciously about me (a skeptic of a company he stands to make bil

The Trump administration relaunches efforts to block state AI laws; Sen. Blackburn is leading negotiations and pushing KOSA as part of an AI preemption package (Axios)

SafetyDGX agent

Axios: The Trump administration relaunches efforts to block state AI laws; Sen. Blackburn is leading negotiations and pushing KOSA as part of an AI preemption package — The White House is negotiating

There is much AI in the news that I initally completely misread this headline 🤣

SafetyDGX agent

Gary Marcus shared a humorous post on X about misreading an AI-related headline due to the current saturation of AI news coverage. The post reflects on how the prevalence of AI stories in media can le

this is gonna go great

SafetyDGX agent

this is gonna go great xAI’s development of artificial intelligence is “a mess”, Co-Executive Editor @mvpeers says. “Elon has fired most of the people who he originally hired at xAI.' 'He has a tenden

Translate-R1: Cost-Aware Translation Tool Use via Reinforcement Learning

SafetyDGX agent

arXiv:2606.06835v1 Announce Type: new Abstract: The performance gap across languages in LLMs is well documented, and closing it natively requires pretraining or fine-tuning on corpora that, for most l

TrioPose: Native Triple-Stream Diffusion Transformers for Pose-Guided Text-to-Image Generation

SafetyDGX agent

arXiv:2606.07053v1 Announce Type: new Abstract: Pose-guided text-to-image generation often suffers from limb distortions and feature crosstalk in complex multi-person scenarios. While existing UNet-ba

UK PM Keir Starmer says tech companies must introduce 'device controls' that stop kids from sending and receiving nude images or face laws forcing them to do so (Reuters)

SafetyDGX agent

Reuters: UK PM Keir Starmer says tech companies must introduce “device controls” that stop kids from sending and receiving nude images or face laws forcing them to do so — Big tech firms operating in

Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows

SafetyDGX agent

arXiv:2606.06875v1 Announce Type: new Abstract: Diffusion transformers (DiTs) equipped with multimodal attention (MM-Attn) have become a dominant paradigm for image generation. However, preventing the

update: the person who posted the original has (rare!) acknowledged the error and delete the original post.

SafetyDGX agent

Gary Marcus posted an update on X noting that the original poster of a viral claim acknowledged their error and deleted the post, highlighting a rare instance of public correction on social media. The

VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models

SafetyDGX agent

arXiv:2602.03160v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

SafetyDGX agent

arXiv:2606.07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal,

What Do People Actually Want From AI? Mapping Preference Plurality

SafetyDGX agent

arXiv:2606.06674v1 Announce Type: new Abstract: Large Language Models (LLMs) are often fine-tuned through Reinforcement Learning from Human Feedback (RLHF) to align with people's preferences and value

What if AI turns out to be a lot less profitable than we have been told? @kirtlenus explores the economic risks https://thecritic.co.uk/the-…

SafetyDGX agent

This article explores potential economic downsides to artificial intelligence, questioning assumptions about AI's profitability and examining financial risks that may have been underestimated or overl

What Is My Robot Thinking? Design Considerations for Transparent and Trustworthy Shared Autonomy

SafetyDGX agent

arXiv:2606.06870v1 Announce Type: new Abstract: Assistive robots operating under shared autonomy must balance user control with autonomous assistance. Because robot actions depend on internal intent i

What Matters When Cotraining Robot Manipulation Policies on Everyday Human Videos?

SafetyDGX agent

arXiv:2606.06627v1 Announce Type: cross Abstract: Human video datasets used for cotraining robot manipulation policies largely consist of curated demonstrations where motions are orchestrated to resem

Where to Touch, How to Contact: Hierarchical RL-MPC Framework for Geometry-Aware Long-Horizon Dexterous Manipulation

SafetyDGX agent

arXiv:2601.10930v3 Announce Type: replace Abstract: A key challenge in contact-rich dexterous manipulation is the need to jointly reason over global geometry and nonsmooth contact dynamics. End-to-end

Which title is better? Seven [Lies/Myths/Bitter Truths] About AI

SafetyDGX agent

Gary Marcus critiques common misconceptions about AI, presenting seven corrective perspectives on prevailing beliefs about artificial intelligence's capabilities and limitations. The post likely addre

Workflow-to-Skill: Skill Creation via Routing-Workflow-Semantics-Attachments Decomposition

SafetyDGX agent

arXiv:2606.06893v1 Announce Type: new Abstract: Large language model agents increasingly rely on Skills to encode procedural knowledge, yet high-quality Skills remain costly to hand-write. This paper

Would you buy SpaceX at the proposed price of $135 a share?

SafetyDGX agent

Gary Marcus poses a hypothetical investment question about SpaceX's valuation at $135 per share, likely exploring perspectives on the company's market value, growth prospects, and investment merit. Th

7 Jun 2026

a riff on https://quoteinvestigator.com/2011/07/09/poker-patsy/

SafetyDGX agent

Gary Marcus likely references the 'poker patsy' concept, commonly attributed to Warren Buffett, which warns that if you've been in a poker game for a while and haven't identified the fool at the table

AI is mentioned more often in SpaceX’s S-1 than Jesus is mentioned in the Bible.

SafetyDGX agent

Gary Marcus compares the frequency of 'AI' mentions in SpaceX's SEC filing (S-1 document) to the frequency of 'Jesus' mentions in the Bible, suggesting that SpaceX emphasizes artificial intelligence e

another take on hard fork and IPOs (again i haven’t listened)

SafetyDGX agent

another take on hard fork and IPOs (again i haven’t listened) I just listened to it. In fairness, they never said anything along the lines of encouraging anyone to buy the IPOs. They also discuss how

as always beware of scammers; here’s a new one, a fake account trying to fool people into thinking i am recommending some garbage. please be…

SafetyDGX agent

Gary Marcus warned X users about a scam involving fake accounts impersonating him to fraudulently promote products or services by falsely claiming his endorsement. The post alerts followers to be caut

blast from the past 3.5 years ago; some things have changed (esp. coding and math, via neurosymbolic techniques) but many haven’t:

SafetyDGX agent

blast from the past 3.5 years ago; some things have changed (esp. coding and math, via neurosymbolic techniques) but many haven’t: Bottom line: From the outset Large Language Models like GPT-3 have gr

California Is Blocking a Federal Audit of Its Voter Rolls California allows first-time voters to register using forms of ID that most Americ…

SafetyDGX agent

California Is Blocking a Federal Audit of Its Voter Rolls California allows first-time voters to register using forms of ID that most Americans would find surprising, including: -Gym membership card -

From fastest growing company to “worst value among its peers” in 18 months. Why, oh why, stick with the CEO?

SafetyDGX agent

From fastest growing company to “worst value among its peers” in 18 months. Why, oh why, stick with the CEO? PitchBook's analysts just ranked OpenAI last for value among its AI peers. Not last for cap

If you aren’t one of the banks running the SpaceX IPO, you’re the mark:

SafetyDGX agent

This post likely critiques the SpaceX IPO process, suggesting that only major financial institutions acting as underwriters benefit substantially from the offering while retail investors and others ar

Imaginary conversations that might actually have happened Act I Sam: We missed all our metrics, Anthropic and Google have gained on us. Give…

SafetyDGX agent

Imaginary conversations that might actually have happened Act I Sam: We missed all our metrics, Anthropic and Google have gained on us. Give us 40 billion dollars. Masa: No way! Sam: If you don't, we

Indeed, I don’t think Trump has thought through the implications.

SafetyDGX agent

Indeed, I don’t think Trump has thought through the implications. This will simply provide even more incentive- as if any was needed - for UK/European players & states to head towards Sovereign AI, ai

Now that the public is waking up to the enormous costs to society of AI, the new dirty trick is to spin the costs of AI as if they were marg…

SafetyDGX agent

Gary Marcus argues that as public awareness grows regarding AI's societal costs, a rhetorical strategy is being employed to reframe or minimize these costs by presenting them as marginal or inevitable

One of the more interesting takes on positive alignment that have recently come out-it’s long and interesting, combining philosophy and trai…

SafetyDGX agent

One of the more interesting takes on positive alignment that have recently come out-it’s long and interesting, combining philosophy and training setups (eg reward proposals), and worth a read. What ha

Only in an America can an industry that has collectively lost over half a trillion dollars —at a pace of roughly a million dollars a minute …

SafetyDGX agent

Gary Marcus critiques the AI industry for its massive financial losses, noting that collectively the sector has lost over half a trillion dollars at an alarming rate of approximately one million dolla

Pro tip: if the IPO you are thinking of investing in is trying to sell shares to the government, it may not be a good sign.

SafetyDGX agent

Gary Marcus cautions that when an IPO company attempts to sell shares to the government as part of its offering, it may indicate underlying weaknesses or lack of confidence from private investors. Thi

Refreshing

SafetyDGX agent

Refreshing What happens when AIs become smarter than us? Why would they keep humans around if given the choice? Our new paper argues that only trying to control AIs is a limited strategy, and that a s

[Revised] Hard Fork’s somewhat soft comments on the IPO, from some who listened

SafetyDGX agent

[Revised] Hard Fork’s somewhat soft comments on the IPO, from some who listened @GaryMarcus at the stage of the Gilded Age II grift cycle when you find out why exchanges put rules in place to protect

the entire field is still shaky on Step 2

SafetyDGX agent

Gary Marcus suggests that Step 2 of some process or framework in AI remains uncertain or unreliable, indicating foundational instability in that particular stage. Without access to the full thread con

← Previous
1…8586878889…214
Next →