AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
All
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,702 results
Safety

The Greedy Advantage in Finite-Horizon Bandits

DGX agent

arXiv:2607.29375v1 Announce Type: cross Abstract: Organizations increasingly rely on sequential experimentation to improve decision-making. While the multi-armed bandit literature has developed algori

safetyarxiv-cs-lg
3 Aug 2026
Safety

The K-Space Signature: Frequency-Domain Representation Learning for Medical Deepfake Detection

Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2607.29541v1 Announce Type: new Abstract: In medical imaging, generative models are increasingly deployed to synthesize realistic data and augment limited datasets. Unfortunately, while benefici

safetyarxiv-cs-cv
3 Aug 2026
Safety

The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations

DGX agent

arXiv:2607.29624v1 Announce Type: cross Abstract: Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback. Conv

safetyarxiv-cs-ai
3 Aug 2026
Safety

There should be a public investigation into the AI hacking incidents by OpenAI and Anthropic. We deserve to know whether these labs are genu…

DGX agent

There should be a public investigation into the AI hacking incidents by OpenAI and Anthropic. We deserve to know whether these labs are genuinely world-class security organizations facing a novel thre

safetygary-marcus--x
3 Aug 2026
Safety

this is the real reason people from OpenAI etc are desperate to shut me up. Astra (which didn’t even have a control group and is maybe not t…

DGX agent

this is the real reason people from OpenAI etc are desperate to shut me up. Astra (which didn’t even have a control group and is maybe not that much better than Fable and certainly not ASI) was perhap

safetygary-marcus--x
3 Aug 2026
Safety

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

DGX agent

arXiv:2607.29254v1 Announce Type: new Abstract: AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential

safetyarxiv-cs-ai
3 Aug 2026
Safety

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking

DGX agent

arXiv:2604.01207v2 Announce Type: replace Abstract: Existing 3D Gaussian Splatting (3DGS) editing methods primarily focus on appearance modification and often struggle to support flexible geometry edi

safetyarxiv-cs-cv
3 Aug 2026
Safety

TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning

DGX agent

arXiv:2607.29586v1 Announce Type: cross Abstract: The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and apply it to a

safetyarxiv-cs-ai
3 Aug 2026
Safety

TRACT: Temporally Routed Action Chunks with Chronological Phase Authority for Contact-Rich Manipulation

DGX agent

arXiv:2607.29285v1 Announce Type: new Abstract: Action chunking shortens the effective decision horizon of robot imitation learning by predicting multiple future actions, while conventional phase cond

safetyarxiv-cs-ro
3 Aug 2026
Safety

TransGraspNet: Physically and Geometrically Consistent Manipulation of Transparent Labware

DGX agent

arXiv:2607.29567v1 Announce Type: new Abstract: Manipulating transparent laboratory glassware that contains liquid is inherently safety-critical: even small geometric errors can cause unstable grasps

safetyarxiv-cs-ro
3 Aug 2026
Safety

Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration

DGX agent

arXiv:2607.28650v1 Announce Type: cross Abstract: While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration of GenAI i

safetyarxiv-cs-ai
3 Aug 2026
Safety

Understanding Alignment in Multimodal LLMs: A Comprehensive Study

DGX agent

Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively under

safetyapple-ml-research
3 Aug 2026
Safety

Unified continuous-time q-learning for mean-field game and mean-field control problems

DGX agent

arXiv:2407.04521v3 Announce Type: replace-cross Abstract: This paper studies the continuous-time q-learning in mean-field jump-diffusion models in a setting where the environment simulator does not pr

safetyarxiv-cs-lg
3 Aug 2026
Safety

VFAD: Variational Semantic Prompting Meets Frequency-Adaptive Representation Learning for Zero-Shot Anomaly Detection

DGX agent

arXiv:2607.29370v1 Announce Type: new Abstract: Zero-shot anomaly detection (ZSAD) aims to detect and localize anomalies in unseen categories without access to target-specific training data. Although

safetyarxiv-cs-cv
3 Aug 2026
Safety

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

DGX agent

arXiv:2607.28678v1 Announce Type: new Abstract: Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally

safetyarxiv-cs-ai
3 Aug 2026
Safety

Vision-based Goal-Reaching Control for Mobile Robots Using a Hierarchical Learning Framework

DGX agent

arXiv:2601.00610v2 Announce Type: replace Abstract: Reinforcement learning (RL) has strong potential in robotics, but exploration-based training complicates safe deployment on large-scale robots. For

safetyarxiv-cs-ro
3 Aug 2026
Safety

WaMo: Wavelet-Enhanced Multi-Frequency Trajectory Analysis for Fine-Grained Text-Motion Retrieval

DGX agent

arXiv:2508.03343v2 Announce Type: replace Abstract: Text-Motion Retrieval (TMR) aims to retrieve 3D motion sequences semantically relevant to text descriptions. However, matching 3D motions with text

safetyarxiv-cs-cv
3 Aug 2026
Safety

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

DGX agent

arXiv:2607.29613v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods,

safetyarxiv-cs-cl
3 Aug 2026
Safety

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

DGX agent

arXiv:2607.29617v1 Announce Type: cross Abstract: Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model

safetyarxiv-cs-ai
3 Aug 2026
Safety

When Unlearning Fails: Reliable Data Deletion under Post-Training in Agent Networks

DGX agent

arXiv:2607.28829v1 Announce Type: cross Abstract: Self-improving federated agent networks keep training after deployment by collecting new trajectories with the current policy and feeding them back in

safetyarxiv-cs-lg
3 Aug 2026
Safety

A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety (Ben Cohen/Wall Street Journal)

DGX agent

Ben Cohen / Wall Street Journal: A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety — Jacob Tsim

safetytechmeme
2 Aug 2026
Safety

Checkmate: you can’t take the harness (which is typically in large part symbolic) away from the neural model without giving up performance. …

DGX agent

Checkmate: you can’t take the harness (which is typically in large part symbolic) away from the neural model without giving up performance. HUGE victory for neurosymbolic AI, straight from @AnthropicA

safetygary-marcus--x
2 Aug 2026
Safety

exactly. math isn’t done. not at all.

DGX agent

exactly. math isn’t done. not at all. I don’t think being critical of the amazing work AI is doing in pure math is fair to @OpenAI until I can start to say why I feel it’s not yet at the level of our

safetygary-marcus--x
2 Aug 2026
Safety

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models …

DGX agent

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models focus on distance while missing that the car itself must rea

safetygary-marcus--x
2 Aug 2026
Safety

OpenAI guy lies about my intent. I would absolutely love to see progress in AI for science and medicine. I have said that here, in my books,…

DGX agent

OpenAI guy lies about my intent. I would absolutely love to see progress in AI for science and medicine. I have said that here, in my books, on countless podcasts, in multiple NYT opeds, in the US Sen

safetygary-marcus--x
2 Aug 2026
Safety

This wins the prize for sleazy misrepresentation. @mattShumer took my 2023 argument for hybridizing LLMs with symbolic tools – which is *exa…

DGX agent

This wins the prize for sleazy misrepresentation. @mattShumer took my 2023 argument for hybridizing LLMs with symbolic tools – which is *exactly* what everyone does nowadays – and made it sound like I

safetygary-marcus--x
2 Aug 2026
Safety

yep they are indeed trying to gaslight me, — in exactly the way you anticipated. both predictable and intellectually dishonest.

DGX agent

yep they are indeed trying to gaslight me, — in exactly the way you anticipated. both predictable and intellectually dishonest. To be clear: when I say “LLM,” I mean the base model, not those with pat

safetygary-marcus--x
2 Aug 2026
Safety

Yet another paper argues that LLMs aren’t close to doing real discovery.

DGX agent

Yet another paper argues that LLMs aren’t close to doing real discovery. MIT and Harvard argue LLMs are nowhere near doing real scientific discovery. They published a paper called “Evaluating Large La

safetygary-marcus--x
2 Aug 2026
Safety

Github repo to learn the OPD/OPSD and how they perform compared to GRPO, on a consumer grade GPU [P]

DGX agent

I am trying to learn concepts like On Policy Distillation (OPD), On Policy Self Distillation (OPSD) and how do they compare to RL algorithms like GRPO. There are a lot of papers on this, but because o

safetyr-machinelearning
1 Aug 2026
Safety

If Leopold had read this on June 26 and trimmed his bets accordingly, SALP would not have melted down. I laid everything out. https://open.s…

DGX agent

If Leopold had read this on June 26 and trimmed his bets accordingly, SALP would not have melted down. I laid everything out. https://open.substack.com/pub/garymarcus/p/the-month-generative-ai-lost-it

safetygary-marcus--x
1 Aug 2026
Safety

The “AGI-is-near” community keeps committing the same logical fallacy over and over; I have seen it at least half a dozen times today alone.…

DGX agent

The “AGI-is-near” community keeps committing the same logical fallacy over and over; I have seen it at least half a dozen times today alone. Every time there’s an advance, I see the same error. Here’s

safetygary-marcus--x
1 Aug 2026
Safety

A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response

DGX agent

arXiv:2607.27597v1 Announce Type: new Abstract: Recent advances in Vision Language Models (VLMs) have created new opportunities for disaster response, where responders must interpret large volumes of

safetyarxiv-cs-ro
31 Jul 2026
Safety

Active Lubrication of Transluminal Medical Instruments

DGX agent

arXiv:2506.07225v2 Announce Type: replace-cross Abstract: Transluminal minimally invasive surgery uses natural orifices and small incisions to access internal anatomical structures, promoting quicker

safetyarxiv-cs-ro
31 Jul 2026
Safety

AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026

DGX agent

arXiv:2607.27228v1 Announce Type: new Abstract: Most conferences rely on peer-review of submissions, but as generative AI makes it easier than ever to prepare submission materials, some conferences ar

safetyarxiv-cs-cl
31 Jul 2026
Safety

AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design Stages

DGX agent

arXiv:2505.10300v2 Announce Type: replace-cross Abstract: Responsible AI (RAI) efforts increasingly emphasize the importance of addressing potential harms early in the AI development lifecycle through

safetyarxiv-cs-ai
31 Jul 2026
Safety

AI Security Priorities: A Field-Wide Agenda

DGX agent

arXiv:2607.26069v1 Announce Type: cross Abstract: As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI securit

safetyarxiv-cs-ai
31 Jul 2026
Safety

Anthropic employee vouches for fiancee of Dario’s chief of staff who holds a lot of Anthropic stock… Where do the Anthropic employees get th…

DGX agent

Anthropic employee vouches for fiancee of Dario’s chief of staff who holds a lot of Anthropic stock… Where do the Anthropic employees get their media training, exactly? Prediction: SALP will be bigger

safetygary-marcus--x
31 Jul 2026
Safety

APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

DGX agent

arXiv:2607.28553v1 Announce Type: new Abstract: Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While flow-matching models (, FlowDPO) h

safetyarxiv-cs-lg
31 Jul 2026
Safety

Ask don't tell: Reducing sycophancy in large language models

DGX agent

arXiv:2602.23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an align

safetyarxiv-cs-ai
31 Jul 2026
Safety

Belief-Guided Decision Making with Uncertainty Gating in the Game of Go

DGX agent

arXiv:2607.26946v1 Announce Type: new Abstract: Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural n

safetyarxiv-cs-ai
31 Jul 2026
Safety

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm

DGX agent

arXiv:2607.27851v1 Announce Type: new Abstract: Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience

safetyarxiv-cs-cl
31 Jul 2026
Safety

BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models

DGX agent

arXiv:2512.00807v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation. Current fairness interve

safetyarxiv-cs-ai
31 Jul 2026
Safety

Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses

DGX agent

arXiv:2607.26639v1 Announce Type: cross Abstract: A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% def

safetyarxiv-cs-ai
31 Jul 2026
Safety

BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

DGX agent

arXiv:2607.27366v1 Announce Type: new Abstract: While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended humanit

safetyarxiv-cs-cl
31 Jul 2026
Safety

Certifying when decision-time information justifies adaptive experimentation

DGX agent

arXiv:2607.27651v1 Announce Type: new Abstract: Adaptive laboratories choose measurements during experiments, yet most methods begin after adaptation is permitted. We introduce Opportunity-aware Polic

safetyarxiv-cs-lg
31 Jul 2026
Safety

Class-Aware Reinforcement Learning for Counterfactual Explanation Generation

DGX agent

arXiv:2607.27905v1 Announce Type: new Abstract: Counterfactual explanations (CFEs) enhance the interpretability of black-box models by generating alternative instances with adjusted feature values tha

safetyarxiv-cs-lg
31 Jul 2026
Safety

Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters

DGX agent

arXiv:2607.27594v1 Announce Type: new Abstract: Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings. However, a

safetyarxiv-cs-lg
31 Jul 2026
Safety

Context-Informed Ship Trajectory Prediction via Conditional Attention

DGX agent

arXiv:2607.27418v1 Announce Type: new Abstract: Long-term ship trajectory prediction is a fundamental capability for maritime safety and autonomous navigation. While recent Transformer-based architect

safetyarxiv-cs-lg
31 Jul 2026
← Previous
1…2425262728…265
Next →