AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,813 results
29 May 2026

Provably Secure Agent Guardrail

SafetyDGX agent

arXiv:2605.29251v1 Announce Type: new Abstract: As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates

Quantifying and Optimizing Simplicity via Polynomial Representations

SafetyDGX agent

arXiv:2605.29823v1 Announce Type: new Abstract: Deep networks often exhibit a preference for 'simple' solutions, and such a simplicity bias is widely believed to play a key role in generalization. Yet

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

SafetyDGX agent

arXiv:2605.28899v1 Announce Type: cross Abstract: Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses si


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

quick, name a prominent company that achieved a major win via tokenmaxxing!

SafetyDGX agent

quick, name a prominent company that achieved a major win via tokenmaxxing! I've yet to see, or hear of a company that is winning against its competition, because said company is spending more on AI t

Quotient DAGs for Off-Policy Evaluation:Forward-Flow Importance Sampling and Exact Slate Propensities

SafetyDGX agent

arXiv:2605.29500v1 Announce Type: cross Abstract: Off-policy evaluation estimates how a target policy would perform using data collected by a different behavior policy, which is crucial when online te

Real-Time Retargeting Using Controllability Boundary for Chandrayaan-3 Lunar Landing

SafetyDGX agent

arXiv:2605.29412v1 Announce Type: cross Abstract: This paper presents the real-time retargeting guidance policy developed for the Chandrayaan-3 lunar landing mission. The baseline guidance generates a

Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers

SafetyDGX agent

arXiv:2601.22139v2 Announce Type: replace-cross Abstract: Reasoning-oriented Large Language Models (LLMs) have achieved remarkable progress with Chain-of-Thought (CoT) prompting, yet they remain funda

ReasonLight: A Multimodal Foundation Model-Enhanced Reinforcement Learning Framework for Zero-Shot Traffic Signal Control

SafetyDGX agent

arXiv:2605.29425v1 Announce Type: new Abstract: Reinforcement learning (RL) has shown promise in traffic signal control (TSC). However, its reliance on predefined states limits responsiveness to obser

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

SafetyDGX agent

arXiv:2605.29447v1 Announce Type: cross Abstract: While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge th

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games

SafetyDGX agent

arXiv:2602.20141v2 Announce Type: replace Abstract: Mean Field Games (MFGs) provide a principled framework for modelling interactions in large population systems. However, algorithmic progress has bee

Reinforcing Few-step Generators via Reward-Tilted Distribution Matching

SafetyDGX agent

arXiv:2605.26108v2 Announce Type: replace Abstract: Recent advances in few-step diffusion distillation have enabled efficient image generation, yet aligning these models with human preferences remains

Replicable Simulation-Based Robot Validation through Provenance

SafetyDGX agent

arXiv:2605.29973v1 Announce Type: new Abstract: Robot behavior is often validated through simulation-based testing, yet the replicability of such campaigns depends critically on transparent documentat

Representation Alignment Rests on Linear Structure

SafetyDGX agent

arXiv:2605.28870v1 Announce Type: cross Abstract: We investigate the Platonic Representation Hypothesis (PRH) through a tripartite statistical framework of representations: signal, bias, and noise. {1

Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents

SafetyDGX agent

arXiv:2605.28850v1 Announce Type: new Abstract: We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments. Using TradeArena, an

Resolution as a Direction: Vector-Panning Feature Alignment for Cross-Resolution Re-Identification

SafetyDGX agent

arXiv:2510.00936v2 Announce Type: replace Abstract: Cross-resolution person re-identification (CR-ReID) remains challenging in practical surveillance, where camera quality and capture distance lead to

Resolving Endpoint Underfitting in Diffusion Bridges via Noise Alignment

SafetyDGX agent

arXiv:2605.28962v1 Announce Type: new Abstract: Diffusion bridge models offer a powerful framework for connecting two data distributions, such as in image restoration and translation. Many existing me

REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image

SafetyDGX agent

arXiv:2605.30338v1 Announce Type: new Abstract: Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation-ready digital assets for applic

Review Arcade: On the Human Alignment and Gameability of LLM Reviews

SafetyDGX agent

arXiv:2605.28897v1 Announce Type: new Abstract: LLM-generated reviews for scientific papers are gaining considerable traction and are even being officially piloted by major conferences. We have to ass

RL2ML: Finite-Rollout Surrogate Objectives from Reinforcement Learning to Maximum Likelihood

SafetyDGX agent

arXiv:2605.30154v1 Announce Type: new Abstract: Correctness-based Reinforcement Learning with Verifiable Rewards (RLVR) trains language models from binary feedback on sampled outputs, but the objectiv

Robust and Generalizable Safety Steering for Text-to-Image Diffusion Transformers

SafetyDGX agent

arXiv:2605.30049v1 Announce Type: new Abstract: Diffusion Transformers have become a powerful backbone for text-to-image generation, but their layered and cross-modal generation process makes safety c

Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training

SafetyDGX agent

arXiv:2603.00454v2 Announce Type: replace-cross Abstract: Generative Flow Networks (GFlowNets) enable fine-tuning large language models to approximate reward-proportional posteriors, but they remain p

RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains

SafetyDGX agent

arXiv:2605.29156v1 Announce Type: cross Abstract: Pointwise reward modeling offers critical signals for LLM post-training, yet struggles with absolute scoring in subjective, non-verifiable settings. R

Rubric-Guided Process Reward for Stepwise Model Routing

SafetyDGX agent

arXiv:2605.29310v1 Announce Type: new Abstract: Stepwise model routing improves the efficiency of Large Reasoning Models (LRMs) by assigning each reasoning step to a suitable model. Recent methods for

SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medication Recommendation

SafetyDGX agent

arXiv:2605.29146v1 Announce Type: cross Abstract: Medication recommendation predicts medications for patient visits, but existing methods still face two key challenges. At the model level, traditional

SAHG: Sector-Anisotropic Hyperbolic Graph Model for Social Bot Detection

SafetyDGX agent

arXiv:2605.30166v1 Announce Type: cross Abstract: LLM-driven social bots can generate fluent, human-like text, reducing the discriminative advantage of content-based detection alone. However, coordina

Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models

SafetyDGX agent

arXiv:2605.30251v1 Announce Type: cross Abstract: Large language models (LLMs) often solve a task when all instructions are given in a single prompt, but fail when the same information is revealed gra

Securing SIM-Assisted Wireless Networks via Quantum Reinforcement Learning

SafetyDGX agent

arXiv:2602.13238v2 Announce Type: replace-cross Abstract: Stacked intelligent metasurfaces (SIMs) have recently emerged as a powerful wave-domain technology that enables multi-stage manipulation of el

Self-Play Reinforcement Learning under Imperfect Information in Big 2

SafetyDGX agent

arXiv:2605.28863v1 Announce Type: cross Abstract: Imperfect-information multiplayer games test whether agents can act under hidden information, sparse rewards, and non-stationary opponents. We study t

SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation

SafetyDGX agent

arXiv:2605.30116v1 Announce Type: new Abstract: Distribution Matching Distillation (DMD) is a widely used paradigm for accelerating inference in few-step video diffusion models. However, DMD-style vid

Shame on you Apple, sending @mcuban’s thoughtful email to spam. Seriously? Glad I have learned not to trust your algorithm.

SafetyDGX agent

Gary Marcus criticizes Apple's email filtering system for incorrectly routing Mark Cuban's message to spam, highlighting a failure in Apple's spam detection algorithm. The post reflects concerns about

SigmaMedStat: Temporal Signal Modeling for ICU False Alarm Reduction

SafetyDGX agent

arXiv:2605.29236v1 Announce Type: new Abstract: Alarm fatigue in intensive care units (ICUs) is a well documented patient safety crisis. Clinical monitors generate 350 or more alarms per patient per d

Statistical Embeddings for Similarity, Retrieval, and Interpretable Alignment of Numeric Tabular Datasets

SafetyDGX agent

arXiv:2605.30289v1 Announce Type: new Abstract: Numeric tabular datasets are the dominant data format in scientific practice, yet large language models lack native mechanisms for representing numeric

talking the party line; they switch, he switches. 🤣

SafetyDGX agent

This post by cognitive scientist Gary Marcus appears to comment humorously on political or ideological conformity, suggesting that people and leaders flip their positions based on party alignment rath

Teaching Values to Machines: Simulating Human-Like Behavior in LLMs

SafetyDGX agent

arXiv:2605.30036v1 Announce Type: new Abstract: Large Language Models (LLMs) demonstrate a remarkable capacity to adopt different personas and roles; however, it remains unclear whether they can manif

The Anatomy of Conversational Scams: A Topic-Based Red Teaming Analysis of Multi-Turn Interactions in LLMs

SafetyDGX agent

arXiv:2601.03134v2 Announce Type: replace Abstract: As LLMs gain persuasive capabilities through extended dialogues, they create new opportunities for studying adversarial conversational behavior in e

The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation

SafetyDGX agent

arXiv:2512.10388v2 Announce Type: replace-cross Abstract: Conventional Sequential Recommender Systems (SRS) typically assign unique hash IDs (HID) to construct item embeddings, which mainly capture co

The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models

SafetyDGX agent

arXiv:2605.29123v1 Announce Type: new Abstract: Masked diffusion language models (MDMs) uniquely support any-order generation, with confidence-based decoding currently serving as the de facto standard

The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane

SafetyDGX agent

arXiv:2605.29082v1 Announce Type: new Abstract: AI agents are increasingly expected to operate as digital employees: accessing enterprise data, making decisions, and taking actions autonomously. But a

The Sample Complexity of Multiclass and Sparse Contextual Bandits

SafetyDGX agent

arXiv:2605.29645v1 Announce Type: cross Abstract: We study contextual bandits in the stochastic i.i.d. setting, where a learner observes contexts drawn from an unknown distribution, selects actions fr

Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning

SafetyDGX agent

arXiv:2605.29032v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably

Think Fast, Talk Smart: Partitioning Deterministic and Neural Computation for Structured Health Text Generation

SafetyDGX agent

arXiv:2605.29652v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly being used to generate health text from structured records such as wearable time series, biomarkers, vital

this is on track to be this year’s “what did Ilya see?” rumor that people love that has nothing to do with reality.

SafetyDGX agent

this is on track to be this year’s “what did Ilya see?” rumor that people love that has nothing to do with reality. Im calling BS on this story. 1. That would be 100,000 employees spending 5k/mo each

Today I learned that @Grimezsz has more courage in her pinky than Roon does in his entire cowardly body. Roon made excuses; Grimes defended …

SafetyDGX agent

Today I learned that @Grimezsz has more courage in her pinky than Roon does in his entire cowardly body. Roon made excuses; Grimes defended her views calmly and respectfully, like grownups should. Ope

Tokenmaxxing is so done.

SafetyDGX agent

Tokenmaxxing is so done. TOKENS OR HUMANS? The new corporate trade off as AI costs balloon 'This is the first time ever that I can remember that technology costs the same as people' - @jainarvind 'Com

Toward AI Systems That Understand Self and Others: A Multi-Phase Inference Framework for Human Cognitive Diversity and World-Model Alignment

SafetyDGX agent

arXiv:2605.29930v1 Announce Type: new Abstract: Mutual misunderstanding in contemporary society does not arise merely because people hold different opinions or values. Even under the same observations

Toward User Preference Alignment in LLM Recommendation via Explicit Context Feedback

SafetyDGX agent

arXiv:2605.29141v1 Announce Type: cross Abstract: Traditional recommender systems (RecSys) primarily infer user preferences from implicit signals (such as clicks, watches, and purchases), often neglec

Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation

SafetyDGX agent

arXiv:2605.29430v1 Announce Type: new Abstract: Automatic speech recognition (ASR) is a core component of human--computer interaction and an increasingly important front-end for LLM-based assistants a

TraceCodec: A Compiler-Backed Neural Codec for Stateful Multi-Flow Network Traffic Traces

SafetyDGX agent

arXiv:2605.29941v1 Announce Type: cross Abstract: Critical networking workflows require high-fidelity packet captures (PCAPs) for testing, security analysis, and protocol validation, not just statisti

TRACER: Persistent Regularization for Robust Multimodal Finetuning

SafetyDGX agent

arXiv:2605.29380v1 Announce Type: cross Abstract: Mainstream strategies for finetuning pretrained multimodal models often degrade out-of-distribution (OOD) robustness, a phenomenon known as catastroph

Train the Agent, Not the Expert: Learning to Harness Heterogeneous Experts for Multi-Turn Visual Reasoning

SafetyDGX agent

arXiv:2605.29894v1 Announce Type: new Abstract: Recent progress in computer vision has produced a wide range of powerful specialized models for detection, segmentation, counting, and other visual task

TriSearch: Learning to Optimize Triangulations via Bistellar Flips

SafetyDGX agent

arXiv:2605.30220v1 Announce Type: new Abstract: We introduce TriSearch, a reinforcement learning framework for optimizing objectives over triangulations of a polytope via bistellar flips. The key idea

Uncertainty Estimation via Hyperspherical Confidence Mapping

SafetyDGX agent

arXiv:2605.05964v2 Announce Type: replace Abstract: Quantifying uncertainty in neural network predictions is essential for high-stakes domains such as autonomous driving, healthcare, and manufacturing

User-Aware Active Knowledge Acquisition for Emotional Support Dialogue

SafetyDGX agent

arXiv:2605.29715v1 Announce Type: new Abstract: Emotional support plays an important role in dialogue systems, and its success depends on adapting to a user's evolving and implicit needs across multi-

V2XCrafter: Learning to Generate Driving Scene Across Agents

SafetyDGX agent

arXiv:2605.29471v1 Announce Type: new Abstract: Collaborative driving systems leverage vehicle-to-everything (V2X) communication for multi-agent collaborative perception to enhance driving safety, yet

ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systems

SafetyDGX agent

arXiv:2602.08567v2 Announce Type: replace-cross Abstract: Multi-agent large language model (LLM) systems increasingly consist of agents that observe and respond to one another's outputs. While value a

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

SafetyDGX agent

arXiv:2605.30117v1 Announce Type: new Abstract: Understanding how Vision-Language-Action (VLA) models transform multimodal knowledge into embodied control remains an open challenge. We present VLA-Tra

Was this priced into the $965 billion valuation?

SafetyDGX agent

Was this priced into the $965 billion valuation? This looks like the beginning of the end for OpenAI and Anthropic. The Chinese AI wave did not just cut prices. It destroyed the entire funding logic b

When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming Loop

SafetyDGX agent

arXiv:2605.29267v1 Announce Type: new Abstract: Foundation models are increasingly trained on synthetic data generated by prior model iterations rather than exclusively on real data. This self-consumi

When and How Long? The Readout-Mediator Angle in Temporal Reasoning

SafetyDGX agent

arXiv:2605.29126v1 Announce Type: cross Abstract: A linear probe can decode a representation almost perfectly and yet be completely irrelevant to how the model uses it. On calendar-date duration reaso

xModel-KD: Cross-modal Knowledge Distillation for 3D Scene Perception using LiDAR

SafetyDGX agent

arXiv:2605.30111v1 Announce Type: cross Abstract: Point cloud segmentation is a fundamental task in 3D scene understanding. Its progress is constrained by the high cost and time required for dense 3D

← Previous
1…107108109110111…214
Next →