AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,702 results
3 Aug 2026

The Greedy Advantage in Finite-Horizon Bandits

SafetyDGX agent

arXiv:2607.29375v1 Announce Type: cross Abstract: Organizations increasingly rely on sequential experimentation to improve decision-making. While the multi-armed bandit literature has developed algori

The K-Space Signature: Frequency-Domain Representation Learning for Medical Deepfake Detection

SafetyDGX agent

arXiv:2607.29541v1 Announce Type: new Abstract: In medical imaging, generative models are increasingly deployed to synthesize realistic data and augment limited datasets. Unfortunately, while benefici

The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations

SafetyDGX agent

arXiv:2607.29624v1 Announce Type: cross Abstract: Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback. Conv


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

There should be a public investigation into the AI hacking incidents by OpenAI and Anthropic. We deserve to know whether these labs are genu…

SafetyDGX agent

There should be a public investigation into the AI hacking incidents by OpenAI and Anthropic. We deserve to know whether these labs are genuinely world-class security organizations facing a novel thre

this is the real reason people from OpenAI etc are desperate to shut me up. Astra (which didn’t even have a control group and is maybe not t…

SafetyDGX agent

this is the real reason people from OpenAI etc are desperate to shut me up. Astra (which didn’t even have a control group and is maybe not that much better than Fable and certainly not ASI) was perhap

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

SafetyDGX agent

arXiv:2607.29254v1 Announce Type: new Abstract: AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking

SafetyDGX agent

arXiv:2604.01207v2 Announce Type: replace Abstract: Existing 3D Gaussian Splatting (3DGS) editing methods primarily focus on appearance modification and often struggle to support flexible geometry edi

TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning

SafetyDGX agent

arXiv:2607.29586v1 Announce Type: cross Abstract: The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and apply it to a

TRACT: Temporally Routed Action Chunks with Chronological Phase Authority for Contact-Rich Manipulation

SafetyDGX agent

arXiv:2607.29285v1 Announce Type: new Abstract: Action chunking shortens the effective decision horizon of robot imitation learning by predicting multiple future actions, while conventional phase cond

TransGraspNet: Physically and Geometrically Consistent Manipulation of Transparent Labware

SafetyDGX agent

arXiv:2607.29567v1 Announce Type: new Abstract: Manipulating transparent laboratory glassware that contains liquid is inherently safety-critical: even small geometric errors can cause unstable grasps

Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration

SafetyDGX agent

arXiv:2607.28650v1 Announce Type: cross Abstract: While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration of GenAI i

Understanding Alignment in Multimodal LLMs: A Comprehensive Study

SafetyDGX agent

Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively under

Unified continuous-time q-learning for mean-field game and mean-field control problems

SafetyDGX agent

arXiv:2407.04521v3 Announce Type: replace-cross Abstract: This paper studies the continuous-time q-learning in mean-field jump-diffusion models in a setting where the environment simulator does not pr

VFAD: Variational Semantic Prompting Meets Frequency-Adaptive Representation Learning for Zero-Shot Anomaly Detection

SafetyDGX agent

arXiv:2607.29370v1 Announce Type: new Abstract: Zero-shot anomaly detection (ZSAD) aims to detect and localize anomalies in unseen categories without access to target-specific training data. Although

ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding

SafetyDGX agent

arXiv:2607.28678v1 Announce Type: new Abstract: Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally

Vision-based Goal-Reaching Control for Mobile Robots Using a Hierarchical Learning Framework

SafetyDGX agent

arXiv:2601.00610v2 Announce Type: replace Abstract: Reinforcement learning (RL) has strong potential in robotics, but exploration-based training complicates safe deployment on large-scale robots. For

WaMo: Wavelet-Enhanced Multi-Frequency Trajectory Analysis for Fine-Grained Text-Motion Retrieval

SafetyDGX agent

arXiv:2508.03343v2 Announce Type: replace Abstract: Text-Motion Retrieval (TMR) aims to retrieve 3D motion sequences semantically relevant to text descriptions. However, matching 3D motions with text

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

SafetyDGX agent

arXiv:2607.29613v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training of Vision-Language-Action (VLA) models has shown strong promise for robotic manipulation. Among RL methods,

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

SafetyDGX agent

arXiv:2607.29617v1 Announce Type: cross Abstract: Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model

When Unlearning Fails: Reliable Data Deletion under Post-Training in Agent Networks

SafetyDGX agent

arXiv:2607.28829v1 Announce Type: cross Abstract: Self-improving federated agent networks keep training after deployment by collecting new trajectories with the current policy and feeding them back in

2 Aug 2026

A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety (Ben Cohen/Wall Street Journal)

SafetyDGX agent

Ben Cohen / Wall Street Journal: A profile of Jacob Tsimerman, who won the Fields Medal last week and is taking a leave from the University of Toronto to join OpenAI and work on AI safety — Jacob Tsim

Checkmate: you can’t take the harness (which is typically in large part symbolic) away from the neural model without giving up performance. …

SafetyDGX agent

Checkmate: you can’t take the harness (which is typically in large part symbolic) away from the neural model without giving up performance. HUGE victory for neurosymbolic AI, straight from @AnthropicA

exactly. math isn’t done. not at all.

SafetyDGX agent

exactly. math isn’t done. not at all. I don’t think being critical of the amazing work AI is doing in pure math is fair to @OpenAI until I can start to say why I feel it’s not yet at the level of our

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models …

SafetyDGX agent

LLMs can know a task is impossible and still optimize it anyway. Ask whether to walk or drive to a car wash 50 meters away, and some models focus on distance while missing that the car itself must rea

OpenAI guy lies about my intent. I would absolutely love to see progress in AI for science and medicine. I have said that here, in my books,…

SafetyDGX agent

OpenAI guy lies about my intent. I would absolutely love to see progress in AI for science and medicine. I have said that here, in my books, on countless podcasts, in multiple NYT opeds, in the US Sen

This wins the prize for sleazy misrepresentation. @mattShumer took my 2023 argument for hybridizing LLMs with symbolic tools – which is *exa…

SafetyDGX agent

This wins the prize for sleazy misrepresentation. @mattShumer took my 2023 argument for hybridizing LLMs with symbolic tools – which is *exactly* what everyone does nowadays – and made it sound like I

yep they are indeed trying to gaslight me, — in exactly the way you anticipated. both predictable and intellectually dishonest.

SafetyDGX agent

yep they are indeed trying to gaslight me, — in exactly the way you anticipated. both predictable and intellectually dishonest. To be clear: when I say “LLM,” I mean the base model, not those with pat

Yet another paper argues that LLMs aren’t close to doing real discovery.

SafetyDGX agent

Yet another paper argues that LLMs aren’t close to doing real discovery. MIT and Harvard argue LLMs are nowhere near doing real scientific discovery. They published a paper called “Evaluating Large La

1 Aug 2026

Github repo to learn the OPD/OPSD and how they perform compared to GRPO, on a consumer grade GPU [P]

SafetyDGX agent

I am trying to learn concepts like On Policy Distillation (OPD), On Policy Self Distillation (OPSD) and how do they compare to RL algorithms like GRPO. There are a lot of papers on this, but because o

If Leopold had read this on June 26 and trimmed his bets accordingly, SALP would not have melted down. I laid everything out. https://open.s…

SafetyDGX agent

If Leopold had read this on June 26 and trimmed his bets accordingly, SALP would not have melted down. I laid everything out. https://open.substack.com/pub/garymarcus/p/the-month-generative-ai-lost-it

The “AGI-is-near” community keeps committing the same logical fallacy over and over; I have seen it at least half a dozen times today alone.…

SafetyDGX agent

The “AGI-is-near” community keeps committing the same logical fallacy over and over; I have seen it at least half a dozen times today alone. Every time there’s an advance, I see the same error. Here’s

31 Jul 2026

A Systems Engineering Framework for Vision-Language-Enabled UAV Triage and Disaster Response

SafetyDGX agent

arXiv:2607.27597v1 Announce Type: new Abstract: Recent advances in Vision Language Models (VLMs) have created new opportunities for disaster response, where responders must interpret large volumes of

Active Lubrication of Transluminal Medical Instruments

SafetyDGX agent

arXiv:2506.07225v2 Announce Type: replace-cross Abstract: Transluminal minimally invasive surgery uses natural orifices and small incisions to access internal anatomical structures, promoting quicker

AI-assisted pre-review of open-source software submissions: an experience report from BOSC 2026

SafetyDGX agent

arXiv:2607.27228v1 Announce Type: new Abstract: Most conferences rely on peer-review of submissions, but as generative AI makes it easier than ever to prepare submission materials, some conferences ar

AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design Stages

SafetyDGX agent

arXiv:2505.10300v2 Announce Type: replace-cross Abstract: Responsible AI (RAI) efforts increasingly emphasize the importance of addressing potential harms early in the AI development lifecycle through

AI Security Priorities: A Field-Wide Agenda

SafetyDGX agent

arXiv:2607.26069v1 Announce Type: cross Abstract: As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI securit

Anthropic employee vouches for fiancee of Dario’s chief of staff who holds a lot of Anthropic stock… Where do the Anthropic employees get th…

SafetyDGX agent

Anthropic employee vouches for fiancee of Dario’s chief of staff who holds a lot of Anthropic stock… Where do the Anthropic employees get their media training, exactly? Prediction: SALP will be bigger

APO: Unsupervised Atomic Policy Optimization for 3D Structure Prediction of Atomic Systems

SafetyDGX agent

arXiv:2607.28553v1 Announce Type: new Abstract: Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While flow-matching models (, FlowDPO) h

Ask don't tell: Reducing sycophancy in large language models

SafetyDGX agent

arXiv:2602.23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an align

Belief-Guided Decision Making with Uncertainty Gating in the Game of Go

SafetyDGX agent

arXiv:2607.26946v1 Announce Type: new Abstract: Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural n

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm

SafetyDGX agent

arXiv:2607.27851v1 Announce Type: new Abstract: Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience

BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models

SafetyDGX agent

arXiv:2512.00807v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation. Current fairness interve

Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses

SafetyDGX agent

arXiv:2607.26639v1 Announce Type: cross Abstract: A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an average 99% def

BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

SafetyDGX agent

arXiv:2607.27366v1 Announce Type: new Abstract: While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended humanit

Certifying when decision-time information justifies adaptive experimentation

SafetyDGX agent

arXiv:2607.27651v1 Announce Type: new Abstract: Adaptive laboratories choose measurements during experiments, yet most methods begin after adaptation is permitted. We introduce Opportunity-aware Polic

Class-Aware Reinforcement Learning for Counterfactual Explanation Generation

SafetyDGX agent

arXiv:2607.27905v1 Announce Type: new Abstract: Counterfactual explanations (CFEs) enhance the interpretability of black-box models by generating alternative instances with adjusted feature values tha

Compliance2LoRA: On-Demand Safety Alignment on Arbitrary Policy Subsets via Hypernetwork-Generated LoRA Adapters

SafetyDGX agent

arXiv:2607.27594v1 Announce Type: new Abstract: Post-training alignment in large reasoning models (LRMs) has significantly improved their adaptability to diverse safety compliance settings. However, a

Context-Informed Ship Trajectory Prediction via Conditional Attention

SafetyDGX agent

arXiv:2607.27418v1 Announce Type: new Abstract: Long-term ship trajectory prediction is a fundamental capability for maritime safety and autonomous navigation. While recent Transformer-based architect

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

SafetyDGX agent

arXiv:2607.28026v1 Announce Type: new Abstract: Recent advances in post-training Large Language Models (LLMs) increasingly rely on Reinforcement Learning with Verifiable Rewards (RLVR) or On-Policy Se

Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response

SafetyDGX agent

arXiv:2607.27508v1 Announce Type: new Abstract: Assistance games formalize human-robot collaboration under asymmetric information: the human knows the goal, while the robot must infer it from observat

Creative Transformation in Literary Texts: Modelling Change Across Representational Levels

SafetyDGX agent

arXiv:2607.28513v1 Announce Type: new Abstract: Creativity is often framed as the production of novelty, yet many cultural works emerge through transformation of earlier artifacts and not through isol

DAS-PMVC: A Framework for Partial Multi-View Clustering via Dual Alignment and Structure Enhancement

SafetyDGX agent

arXiv:2607.27761v1 Announce Type: cross Abstract: In recent years, multi-view clustering has attracted widespread research interest. However, due to limitations in data collection devices, data across

DECODE: Tackling Representation and Decision Degradation in Continual AI-Generated Image Detection

SafetyDGX agent

arXiv:2607.27882v1 Announce Type: new Abstract: As generative models continue to evolve, AI-generated image detectors must incrementally adapt to emerging generative domains while preserving knowledge

DexDirect: Direct Kinesthetic Arm Guidance for Efficient Dexterous Demonstration Collection

SafetyDGX agent

arXiv:2607.27784v1 Announce Type: new Abstract: Scalable collection of dexterous manipulation demonstrations remains a major bottleneck for robot learning. High-fidelity interfaces often require costl

Digital Harf: A Clinically Integrated Multimodal AI System for Pervasive Arabic Speech and Language Therapy

SafetyDGX agent

arXiv:2607.27212v1 Announce Type: cross Abstract: Children with Autism Spectrum Disorder in Arabic-speaking countries face compounded barriers to effective speech and language therapy: a shortage of q

DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation

SafetyDGX agent

arXiv:2607.27614v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-langua

Dynamic Spectral Filtering for Temporal Graph Learning: Learning Evolving Propagation Operators

SafetyDGX agent

arXiv:2607.27891v1 Announce Type: cross Abstract: Temporal graph learning is commonly organized around the evolution of node states or the encoding of interaction histories. We study an underexplored,

Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models

SafetyDGX agent

arXiv:2607.26588v1 Announce Type: new Abstract: The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

SafetyDGX agent

arXiv:2607.28243v1 Announce Type: new Abstract: Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodi

EmoFeedback^2: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback

SafetyDGX agent

arXiv:2511.19982v3 Announce Type: replace Abstract: Continuous emotional image content generation (C-EICG) is emerging rapidly due to its ability to produce images aligned with both user descriptions

← Previous
1…1920212223…212
Next →