AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,356 results
Safety

Policy Improvement Reinforcement Learning

DGX agent

arXiv:2604.00860v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a central post-training paradigm for improving the reasoning capabilities of large

safetyarxiv-cs-lg
29 Apr 2026
Safety

Progressing beyond Art Masterpieces or Touristic Cliches: how to assess your LLMs for cultural alignment?

X Post
Paper
YouTube
Reddit
GitHub
Clear filters
DGX agent

arXiv:2604.25654v1 Announce Type: new Abstract: Although the cultural (mis)alignment of Large Language Models (LLMs) has attracted increasing attention -- often framed in terms of cultural bias -- unt

safetyarxiv-cs-cl
29 Apr 2026
Safety

Records Sam Altman might set: • Most money burned • Biggest lead squandered • Most promises broken • Most nonprofits seized (tie, probably n…

DGX agent

Records Sam Altman might set: • Most money burned • Biggest lead squandered • Most promises broken • Most nonprofits seized (tie, probably nobody has two) • Most times fired from one job (anyone know

safetygary-marcus--x
29 Apr 2026
Safety

Reference-Augmented Learning for Precise Tracking Policy of Tendon-Driven Continuum Robots

DGX agent

arXiv:2604.25698v1 Announce Type: new Abstract: Tendon-Driven Continuum Robots (TDCRs) pose significant control challenges due to their highly nonlinear, path-dependent dynamics and non-Markovian char

safetyarxiv-cs-ro
29 Apr 2026
Safety

Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models

DGX agent

arXiv:2604.25636v1 Announce Type: new Abstract: Unified multimodal models (UMMs) integrate visual understanding and generation within a single framework. For text-to-image (T2I) tasks, this unified ca

safetyarxiv-cs-cv
29 Apr 2026
Safety

ResetEdit: Precise Text-guided Editing of Generated Image via Resettable Starting Latent

DGX agent

arXiv:2604.25128v1 Announce Type: new Abstract: Recent advances in diffusion models have enabled high-quality image generation, leading to increasing demand for post-generation editing that modifies l

safetyarxiv-cs-cv
29 Apr 2026
Safety

ReSim: Reliable World Simulation for Autonomous Driving

DGX agent

arXiv:2506.09981v2 Announce Type: replace Abstract: How can we reliably simulate future driving scenarios under a wide range of ego driving behaviors? Recent driving world models, developed exclusivel

safetyarxiv-cs-cv
29 Apr 2026
Safety

Resource Orchestration & Redistribution Systems As production becomes more efficient (energy, goods, services), the constraint shifts to all…

DGX agent

Resource Orchestration & Redistribution Systems As production becomes more efficient (energy, goods, services), the constraint shifts to allocation. These systems dynamically route resources based on

safetyyohei-nakajima--x
29 Apr 2026
Safety

Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective

DGX agent

arXiv:2510.10150v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) serves as a cornerstone technique for enhancing the reasoning capabilities of Large Language M

safetyarxiv-cs-lg
29 Apr 2026
Safety

“roughly $1.35 trillion of pure narrative” is the best euphemism for pure bullshit that I have ever seen, @gnoble79. 👏

DGX agent

“roughly 1.35 trillion of pure narrative” is the best euphemism for pure bullshit that I have ever seen, @gnoble79. 👏 This is the most OUTRAGEOUS deal I've seen in my 45 years on Wall Street. SpaceX j

safetygary-marcus--x
29 Apr 2026
Safety

Sam’s not the right CEO for OpenAI anymore. He squandered their lead, failed to deliver the AGI he promised, flailed productwise, and torche…

DGX agent

Sam’s not the right CEO for OpenAI anymore. He squandered their lead, failed to deliver the AGI he promised, flailed productwise, and torched his own reputation; meanwhile he is burning money at a unp

safetygary-marcus--x
29 Apr 2026
Safety

Sheer insanity. Amazon, Google, Microsoft, and Meta collectively are spending more money than the Manhattan Project *every single month*. Mo…

DGX agent

Sheer insanity. Amazon, Google, Microsoft, and Meta collectively are spending more money than the Manhattan Project *every single month*. More than 12x the Manhattan Project every year. And what they

safetygary-marcus--x
29 Apr 2026
Safety

Spark Policy Toolkit: Semantic Contracts and Scalable Execution for Policy Learning in Spark

DGX agent

arXiv:2604.25061v1 Announce Type: cross Abstract: Custom policy-learning pipelines in Spark fail for two coupled systems reasons: rowwise Python execution makes inference impractical, and driver-side

safetyarxiv-cs-lg
29 Apr 2026
Safety

Spectral bandits

DGX agent

arXiv:2604.25272v1 Announce Type: cross Abstract: Smooth functions on graphs have wide applications in manifold and semi-supervised learning. In this work, we study a bandit problem where the payoffs

safetyarxiv-cs-lg
29 Apr 2026
Safety

Subliminal Steering: Stronger Encoding of Hidden Signals

DGX agent

arXiv:2604.25783v1 Announce Type: new Abstract: Subliminal learning describes a student language model inheriting a behavioral bias by fine-tuning on seemingly innocuous data generated by a biased tea

safetyarxiv-cs-cl
29 Apr 2026
Safety

the ability of this man to sell things he doesn’t believe in is truly extraordinary

DGX agent

the ability of this man to sell things he doesn’t believe in is truly extraordinary SAM ALTMAN “OpenAI is structured as a nonprofit because we don’t ever want to be making decisions to benefit shareho

safetygary-marcus--x
29 Apr 2026
Safety

The Musk-Altman trial is clearly a battle of egos—and it’s hard to root for either side—𝗯𝘂𝘁 𝗶𝘁’𝘀 𝗮𝗹𝘀𝗼 𝗮 𝘁𝗿𝗶𝗮𝗹 𝗮𝗯𝗼𝘂𝘁 𝘄…

DGX agent

The Musk-Altman trial is clearly a battle of egos—and it’s hard to root for either side—𝗯𝘂𝘁 𝗶𝘁’𝘀 𝗮𝗹𝘀𝗼 𝗮 𝘁𝗿𝗶𝗮𝗹 𝗮𝗯𝗼𝘂𝘁 𝘄𝗵𝗲𝘁𝗵𝗲𝗿 𝗢𝗽𝗲𝗻𝗔𝗜 𝘀𝗵𝗼𝘂𝗹𝗱 𝗯𝗲 𝗵𝗲𝗹𝗱 𝘁𝗼 𝗶𝘁𝘀 𝗽𝗿𝗼𝗺𝗶𝘀𝗲𝘀 𝘁𝗼 𝗯𝗲 𝗮 𝗻𝗼𝗻𝗽𝗿𝗼𝗳𝗶𝘁 𝘄𝗼𝗿𝗸𝗶𝗻𝗴 𝗳𝗼𝗿 𝘁𝗵𝗲 𝗯𝗲𝗻𝗲

safetygary-marcus--x
29 Apr 2026
Safety

The President of the United States trying to take $10 billion dollars from US taxpayers for his own personal gain—and he is probably going t…

DGX agent

This post appears to discuss allegations that a U.S. President sought to obtain $10 billion in taxpayer funds for personal use. The content is incomplete in the provided title, making it difficult to

safetygary-marcus--x
29 Apr 2026
Safety

The Rashomon Effect for Visualizing High-Dimensional Data

DGX agent

arXiv:2604.00485v2 Announce Type: replace Abstract: Dimension reduction (DR) is inherently non-unique: multiple embeddings can preserve the structure of high-dimensional data equally well while differ

safetyarxiv-cs-lg
29 Apr 2026
Safety

The Russian Legislative Corpus

DGX agent

arXiv:2406.04855v3 Announce Type: replace Abstract: We present a comprehensive corpus of Russian primary and secondary legislation adopted between 1991 and 2025, comprising 304,382 texts (194,425,905

safetyarxiv-cs-cl
29 Apr 2026
Safety

Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models

DGX agent

arXiv:2510.16340v2 Announce Type: replace Abstract: Recent advances in post-training techniques have endowed Large Language Models (LLMs) with enhanced capabilities for tackling complex, logic-intensi

safetyarxiv-cs-cl
29 Apr 2026
Safety

this kind of armchair nonsense is especially hilarious on day when I defended Elon’s lawsuit on CNN and in my newsletter. 🙄

DGX agent

this kind of armchair nonsense is especially hilarious on day when I defended Elon’s lawsuit on CNN and in my newsletter. 🙄 @GaryMarcus @gnoble79 I'm not sure history is going to be kind to this take.

safetygary-marcus--x
29 Apr 2026
Safety

Three Models of RLHF Annotation: Extension, Evidence, and Authority

DGX agent

arXiv:2604.25895v1 Announce Type: cross Abstract: Preference-based alignment methods, most prominently Reinforcement Learning with Human Feedback (RLHF), use the judgments of human annotators to shape

safetyarxiv-cs-cl
29 Apr 2026
Safety

To be a bit more clear for people who did not read the full article: A domestic data-center ban does not directly address the risks of extin…

DGX agent

To be a bit more clear for people who did not read the full article: A domestic data-center ban does not directly address the risks of extinction from ASI, and weakens the US. People may or may not al

safetyconnor-leahy--x
29 Apr 2026
Safety

TouchAI: Exploring human-AI perceptual alignment in touch through language model representations

DGX agent

arXiv:2406.06587v2 Announce Type: replace Abstract: Aligning large language models (LLMs) behaviour with human intent is critical for future AI. An important yet often overlooked aspect of this alignm

safetyarxiv-cs-cl
29 Apr 2026
Safety

True or False: OpenAI will eventually become a massively profitable company, more than earning out all the money that went into it.

DGX agent

Gary Marcus poses a question about whether OpenAI will achieve sufficient profitability to justify the substantial capital investments made into the company. This reflects broader industry speculation

safetygary-marcus--x
29 Apr 2026
Safety

Unrequited Emotions: Investigating the Gaps in Motivation and Practice in Speech Emotion Recognition Research

DGX agent

arXiv:2604.25776v1 Announce Type: new Abstract: Critical analyses of emotion recognition technology have raised ethical concerns around task validity and potential downstream impacts, urging researche

safetyarxiv-cs-cl
29 Apr 2026
Safety

Voice, Bias, and Coreference: An Interpretability Study of Gender in Speech Translation

DGX agent

arXiv:2511.21517v2 Announce Type: replace Abstract: Unlike text, speech conveys information about the speaker, such as gender, through acoustic cues like pitch. This gives rise to modality-specific bi

safetyarxiv-cs-cl
29 Apr 2026
Safety

We at @ControlAI are sometimes asked what we think of banning datacenter construction. At CAI, we focus on one issue: The risk of extinction…

DGX agent

We at @ControlAI are sometimes asked what we think of banning datacenter construction. At CAI, we focus on one issue: The risk of extinction from superintelligent AI. The only way to prevent this is t

safetyconnor-leahy--x
29 Apr 2026
Safety

“we might as well stop training radiologists” Geoff Hinton, 2016, vs the actual data, via Torsten Slok at Apollo

DGX agent

Gary Marcus shares Torsten Slok's analysis comparing Geoffrey Hinton's 2016 prediction that radiologist training should cease due to AI capabilities against actual empirical data on AI performance in

safetygary-marcus--x
29 Apr 2026
Safety

When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient

DGX agent

arXiv:2604.25872v1 Announce Type: new Abstract: Training language models via reinforcement learning often relies on imperfect proxy rewards, since ground truth rewards that precisely define the intend

safetyarxiv-cs-lg
29 Apr 2026
Safety

Zuckerberg bets big on biology? About as much he put into 5 employees in his AI startup for a year 🙄 He is *far* more interested AI for sel…

DGX agent

Gary Marcus critiques Mark Zuckerberg's claimed focus on biology research, suggesting the financial commitment is minimal compared to his AI investments and stating that Zuckerberg's actual priorities

safetygary-marcus--x
29 Apr 2026
Model Releases

A BERTology View of LLM Orchestrations: Token- and Layer-Selective Probes for Efficient Single-Pass Classification

DGX agent

arXiv:2601.13288v2 Announce Type: replace Abstract: Production LLM systems often rely on separate models for safety and other classification-heavy steps, increasing latency, VRAM footprint, and operat

model-releasesarxiv-cs-cl
28 Apr 2026
Safety

A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs

DGX agent

arXiv:2603.07475v2 Announce Type: replace Abstract: Autoregressive (AR) language models build representations incrementally via left-to-right prediction, while diffusion language models (dLLMs) are tr

safetyarxiv-cs-cl
28 Apr 2026
Safety

A Differentiable Framework for Global Circulation Model Precipitation Bias Correction

DGX agent

arXiv:2604.23045v1 Announce Type: new Abstract: Systematic biases in Global Circulation Model (GCM) outputs limit their direct applicability in regional planning, necessitating bias correction. Correc

safetyarxiv-cs-lg
28 Apr 2026
Safety

A Multi-Dimensional Audit of Politically Aligned Large Language Models

DGX agent

arXiv:2604.24429v1 Announce Type: new Abstract: As the application of Large Language Models (LLMs) spreads across various industries, there are increasing concerns about the potential for their misuse

safetyarxiv-cs-cl
28 Apr 2026
Safety

A Reward-Free Viewpoint on Multi-Objective Reinforcement Learning

DGX agent

arXiv:2604.24532v1 Announce Type: new Abstract: Many sequential decision-making tasks involve optimizing multiple conflicting objectives, requiring policies that adapt to different user preferences. I

safetyarxiv-cs-lg
28 Apr 2026
Safety

A Taxonomy and Resolution Strategy for Client-Level Disagreements in Federated Learning

DGX agent

arXiv:2604.23386v1 Announce Type: cross Abstract: Federated Learning (FL) typically assumes unconditional collaboration, a premise that overlooks the complexities of real-world, multi-stakeholder envi

safetyarxiv-cs-ai
28 Apr 2026
Safety

AdaRubric: Task-Adaptive Rubrics for LLM Agent Evaluation

DGX agent

arXiv:2603.21362v2 Announce Type: replace Abstract: LLM-as-Judge evaluation fails agent tasks because a fixed rubric cannot capture what matters for this task: code debugging demands Correctness and E

safetyarxiv-cs-ai
28 Apr 2026
Safety

Additive Control Variates Dominate Self-Normalisation in Off-Policy Evaluation

DGX agent

arXiv:2602.14914v2 Announce Type: replace Abstract: Off-policy evaluation (OPE) is essential for assessing ranking and recommendation systems without costly online interventions. Self-Normalised Inver

safetyarxiv-cs-lg
28 Apr 2026
Safety

Adversary-Free Counterfactual Prediction via Information-Regularized Representations

DGX agent

arXiv:2510.15479v2 Announce Type: replace Abstract: We study counterfactual prediction under assignment bias and propose a mathematically grounded, information-theoretic approach that removes treatmen

safetyarxiv-cs-lg
28 Apr 2026
Safety

Algorithmic Administration and the EU AI Act: Legal Principles for Public Sector Use of AI

DGX agent

arXiv:2604.22765v1 Announce Type: cross Abstract: The increasing use of artificial intelligence (AI) by public authorities introduces both opportunities for innovation and significant challenges for t

safetyarxiv-cs-ai
28 Apr 2026
Safety

Aligning with Your Own Voice: Self-Corrected Preference Learning for Hallucination Mitigation in LVLMs

DGX agent

arXiv:2604.24395v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) frequently suffer from hallucinations. Existing preference learning-based approaches largely rely on proprietary mo

safetyarxiv-cs-ai
28 Apr 2026
Safety

ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis

DGX agent

arXiv:2404.10141v2 Announce Type: replace-cross Abstract: Text-to-image (T2I) models have achieved remarkable progress in high-quality image synthesis, yet most benchmarks rely on simple, self-contain

safetyarxiv-cs-cl
28 Apr 2026
Safety

AnemiaVision: Non-Invasive Anemia Detection via Smartphone Imagery Using EfficientNet-B3 with TrivialAugmentWide, Mixup Augmentation, and Persistent Patient History Management

DGX agent

arXiv:2604.22964v1 Announce Type: new Abstract: Anemia affects over one billion people globally and remains severely under-diagnosed in low-resource regions where laboratory blood tests are inaccessib

safetyarxiv-cs-cv
28 Apr 2026
Safety

Animalbooth: multimodal feature enhancement for animal subject personalization

DGX agent

arXiv:2509.16702v2 Announce Type: replace Abstract: Personalized animal image generation is challenging due to rich appearance cues and large morphological variability. Existing approaches often exhib

safetyarxiv-cs-cv
28 Apr 2026
Safety

As an avid cyclist, I was amused to see ChatGPT's “powerful new image engine' draw a bicycle with the 'brake' label pointing to empty space …

DGX agent

As an avid cyclist, I was amused to see ChatGPT's “powerful new image engine' draw a bicycle with the 'brake' label pointing to empty space where brakes are sometimes found on other bicycles. The poin

safetygary-marcus--x
28 Apr 2026
Safety

Bellman Residual Minimization for Control: Geometry, Stationarity, and Convergence

DGX agent

arXiv:2601.18840v3 Announce Type: replace Abstract: Markov decision problems are most commonly solved via dynamic programming. Another approach is Bellman residual minimization, which directly minimiz

safetyarxiv-cs-lg
28 Apr 2026
← Previous
1…241242243244245…300
Next →