AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

safety

GridTimelineEvolution
12,704 results
27 Jul 2026

Quantifying Political Partisanship for Cross-Platform Analyses

SafetyDGX agent

arXiv:2607.21842v1 Announce Type: cross Abstract: Research on political polarization on social media depends on the ability to reliably measure partisanship in user-generated content. However, existin

Reliability Scales Inversely: Bigger Language Models Compound Mistakes Faster

SafetyDGX agent

arXiv:2607.18292v2 Announce Type: replace-cross Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account --

Safe Learning Predictive Control for Ego-World Robotic Systems

SafetyDGX agent

arXiv:2607.22225v1 Announce Type: new Abstract: Safe autonomous navigation in shared environments requires the ability to anticipate and react to the latent behaviors of surrounding robots. In this pa


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Security Without Detection: Economic Denial as a Primitive for Edge and IoT Defense

SafetyDGX agent

arXiv:2512.23849v2 Announce Type: replace-cross Abstract: Sophisticated attackers can evade detection-based security by using encryption, stealth tactics, and low-rate attack patterns. This challenge

Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning

SafetyDGX agent

arXiv:2507.01551v3 Announce Type: replace-cross Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~

Spectral Prior for Reducing Exposure Bias in Diffusion Models

SafetyDGX agent

arXiv:2607.22091v1 Announce Type: new Abstract: Diffusion models typically suffer from error accumulation during iterative sampling, commonly referred to as exposure bias. We reveal systematic frequen

Stocks should be ripping today given the move in oil. But they're not thanks to NVDA and it's unbelievable 250 Billion roundtrip with Open…

SafetyDGX agent

Stocks should be ripping today given the move in oil. But they're not thanks to NVDA and it's unbelievable 250 Billion roundtrip with OpenAI. There's no hiding anymore that the whole AI bubble is real

Tech industry leaders join to form Open Secure AI Alliance to promote safety and security

SafetyDGX agent

Nvidia Corp. today announced the launch of the Open Secure AI Alliance, a new organization founded by technology, cloud computing and cybersecurity leaders to build and share open artificial intellige

TextSLIP: Text Self-Supervised CLIP for Medical Report Generation

SafetyDGX agent

arXiv:2607.21970v1 Announce Type: new Abstract: Automating radiology report generation is important for improving reporting consistency and clinical workflows . While Contrastive Language--Image Pretr

The Coordination Gap: Multi-Agent Alternation Metrics for Temporal Fairness in Repeated Games

SafetyDGX agent

arXiv:2603.05789v5 Announce Type: replace-cross Abstract: Repeated multi-agent interactions require evaluation metrics that capture not only payoff distributions but also their temporal organization.

TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI

SafetyDGX agent

arXiv:2607.22465v1 Announce Type: cross Abstract: Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Ex

Twins: Learn to Predict Unified Representations with Focal Loss

SafetyDGX agent

arXiv:2607.22531v1 Announce Type: new Abstract: Unified multimodal models seek a shared visual token space that supports both multimodal understanding and image generation. Discrete methods unify the

Very cool paper from Microsoft. The idea is to train agents on replayed teacher trajectories instead of live environment rollouts. On-policy…

SafetyDGX agent

Very cool paper from Microsoft. The idea is to train agents on replayed teacher trajectories instead of live environment rollouts. On-policy distillation for agentic tasks is expensive because every u

ViTacWorld: Scaling Visuo-Tactile World Models for Contact-Rich Robot Manipulation

SafetyDGX agent

arXiv:2607.22530v1 Announce Type: new Abstract: Contact-rich robot manipulation requires physical interaction cues that are often invisible to cameras, making tactile sensing essential for robust cont

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

SafetyDGX agent

arXiv:2505.19212v2 Announce Type: replace Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical align

Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

SafetyDGX agent

arXiv:2607.22218v1 Announce Type: new Abstract: Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, rais

26 Jul 2026

important thread; don’t read just the first tweet (which has a caveat in the second).

SafetyDGX agent

important thread; don’t read just the first tweet (which has a caveat in the second). In case you think the problem of science slop is hypothetical: Here's the president of OpenAI retweeting wrong sci

training to benchmarks ≠ getting to AGI

SafetyDGX agent

training to benchmarks ≠ getting to AGI This suggests that Opus 5 's gain on ARC-AGI-3 was the result of specific training to improve on that eval, and not a generalized increase in abstract reasoning

25 Jul 2026

All anyone has to do is go back 15 years and look at all the ridiculous predictions Elmo has made about Mars, self driving cars, his ridicul…

SafetyDGX agent

All anyone has to do is go back 15 years and look at all the ridiculous predictions Elmo has made about Mars, self driving cars, his ridiculous Boring Company. No one should take anything he says seri

but they won’t.

SafetyDGX agent

but they won’t. AI safety experts say OpenAI’s rogue models may mean the company has already blown past its own internal red lines. That would mean it should pause development until it creates better

Cohere has proudly signed on to this letter. The importance of open-source models to the AI ecosystem cannot be understated. We believe ever…

SafetyDGX agent

Cohere has proudly signed on to this letter. The importance of open-source models to the AI ecosystem cannot be understated. We believe everyone, in every country, should have control over the technol

I am surprised how few people are aware that the reasoning for OpenAI/Anthropic models is all encrypted. The 'reasoning' you see in the UI i…

SafetyDGX agent

OpenAI and Anthropic’s language models keep their internal reasoning encrypted; what users see in the UI is only a filtered summary of that reasoning. This practice was highlighted in a tweet by Sarah

The two giant generative AI startups were treated like gods for a couple years. Both are now facing massive pushback.

SafetyDGX agent

The two giant generative AI startups were treated like gods for a couple years. Both are now facing massive pushback. Anthropic employees go nuclear on the popularity of open source AI. It is getting

We recognize there are a lot of questions and speculative details circulating related to the Hugging Face incident. This is an unprecedented…

SafetyDGX agent

We recognize there are a lot of questions and speculative details circulating related to the Hugging Face incident. This is an unprecedented incident, and we think it marks an important moment for AI

When a normal company does a bad thing, they take responsibility + apologize + outline how it won't happen again. OpenAI instead is like 'we…

SafetyDGX agent

When a normal company does a bad thing, they take responsibility + apologize + outline how it won't happen again. OpenAI instead is like 'we are entering a new era. this will happen again. no one amon

Wonderful to have a broad base acknowledgement of the need for open weights. We are delighted to support as signatories. Open weights is the…

SafetyDGX agent

Wonderful to have a broad base acknowledgement of the need for open weights. We are delighted to support as signatories. Open weights is the key to ensuring security, research, innovation and competit

24 Jul 2026

A Framework for Reputation Aware Uninorm-driven Consensus Algorithms for Blockchain Networks

SafetyDGX agent

arXiv:2607.20700v1 Announce Type: cross Abstract: The operation of blockchain is governed by consensus algorithms (CA). Several consensus mechanisms require significant computational power, while othe

A Real-Time Generalized Nash Equilibrium Framework for Interaction-Aware Autonomous Driving in Mixed Traffic

SafetyDGX agent

arXiv:2607.21043v1 Announce Type: new Abstract: Safe and efficient navigation in mixed-traffic environments remains a critical challenge for Autonomous Vehicles (AVs), primarily due to the complex int

A Self-Evolving Default Action for Cooperative Tasks with Continuous Action Space

SafetyDGX agent

arXiv:2607.18597v2 Announce Type: replace Abstract: Counterfactual credit assignment has proven effective in multi-agent reinforcement learning (MARL) for discrete action spaces, yet its extension to

A Unified Moral-Value Dataset for Instruction Tuning

SafetyDGX agent

arXiv:2607.21279v1 Announce Type: new Abstract: Large language models (LLMs) have developed rapidly and become valuable tools in everyday life. However, how to align LLMs to a particular set of human

Adaptive Confidence-weighted Expansion for Trustworthy Multi-Omics Multimodal Fusion

SafetyDGX agent

arXiv:2607.20742v1 Announce Type: new Abstract: Multimodal learning is a robust approach to improve predictive performance in applications such as medical prognosis. However, the clinical applicabilit

Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs

SafetyDGX agent

arXiv:2607.21291v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong generation and reasoning performance, but the Transformer architecture incurs high inference cost. Existing

Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation

SafetyDGX agent

arXiv:2607.18042v2 Announce Type: replace-cross Abstract: End-to-end vision-language navigation (VLN) with causal vision-language models maps instructions and egocentric observations directly to actio

Approximate Quantum State Preparation Through Proximal Policy Optimization

SafetyDGX agent

arXiv:2607.21121v1 Announce Type: cross Abstract: In this work, a quantum architecture search framework for approximate quantum state preparation (QSP) is proposed. QSP is a challenging task, since th

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents

SafetyDGX agent

arXiv:2606.21262v2 Announce Type: replace Abstract: Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or

ASTRA-Net: Anatomy-Specific Transfer and Representation Alignment for Drug-Induced Sleep Endoscopy Segmentation

SafetyDGX agent

arXiv:2607.21370v1 Announce Type: new Abstract: Quantitative drug-induced sleep endoscopy (DISE) requires reliable airway boundaries at specific anatomical levels. Pixel-level DISE annotations are sca

AttriMem: Attribution-Guided Process Feedback for Agent Memory Learning

SafetyDGX agent

arXiv:2607.21106v1 Announce Type: new Abstract: Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what information t

Axolotl3D: a Unified Framework for Faithful 3D Shape Completion

SafetyDGX agent

arXiv:2607.20660v1 Announce Type: new Abstract: Recent 3D generative models produce high-quality geometry from a single image using large-scale priors and diffusion architectures. However, they assume

Belief Propagation in LLM World Models: Measuring Strategic Information Bias with Prediction Markets

SafetyDGX agent

arXiv:2607.20441v1 Announce Type: new Abstract: Every information ecosystem produces beliefs that shape strategic decisions. Both human analysts and AI systems inherit the blind spots of their informa

Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning

SafetyDGX agent

arXiv:2607.21558v1 Announce Type: new Abstract: Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducing sycophancy

Climate-resilient electric vehicle charging infrastructure for sustainable cities: An interpretable causal-ensemble framework for preventive maintenance and low-carbon mobility

SafetyDGX agent

arXiv:2607.21444v1 Announce Type: cross Abstract: Reliable electric vehicle (EV) charging infrastructure is a cornerstone of sustainable, low-carbon cities, yet urban climate stress such as extreme he

Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents

SafetyDGX agent

arXiv:2607.11346v3 Announce Type: replace Abstract: Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable SOP const

Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception

SafetyDGX agent

arXiv:2607.20444v1 Announce Type: cross Abstract: Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal.

CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation

SafetyDGX agent

arXiv:2607.20862v1 Announce Type: new Abstract: At present, reliable evaluation of non-verifiable tasks remains challenging. Existing approaches often fail to adequately capture the diverse evaluative

Detectors Learn the Wrong Thing: Shortcut-Resistant Adversarial Training Against Physically Realizable Attacks

SafetyDGX agent

arXiv:2607.21243v1 Announce Type: new Abstract: AI-enabled visual perception systems are increasingly deployed in intelligent transportation infrastructure and autonomous vehicle related applications.

Diagnosing Pathological Chain-of-Thought in Reasoning Models

SafetyDGX agent

arXiv:2602.13904v2 Announce Type: replace Abstract: Chain-of-thought (CoT) reasoning is fundamental to modern LLM architectures and represents a critical intervention point for AI safety. However, CoT

DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation

SafetyDGX agent

arXiv:2607.21371v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories. While the self-supervised mo

Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval

SafetyDGX agent

arXiv:2607.20984v1 Announce Type: new Abstract: This paper proposes the Distribution-Alignment Bridge (DAB), a framework that reconceptualizes text-to-video retrieval as a distribution alignment task

DTIF: Robust Loop Closure Detection via Delaunay Triangle Topology in Complex Forests

SafetyDGX agent

arXiv:2607.21138v1 Announce Type: new Abstract: Accurate forest inventory and large-scale mapping are essential for ecosystem monitoring and sustainable forest management. Multiple low-cost edge platf

DynaMark: A Reinforcement Learning Framework for Dynamic Watermarking in Industrial Machine Tool Controllers

SafetyDGX agent

arXiv:2508.21797v2 Announce Type: replace-cross Abstract: Industry 4.0's highly networked Machine Tool Controllers (MTCs) are prime targets for replay attacks that use outdated sensor data to manipula

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization

SafetyDGX agent

arXiv:2607.21013v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER to a new leve

EmoSpace: Immersive Affective Image Generation Guided by Fine-Grained Emotion Prototypes

SafetyDGX agent

arXiv:2602.11658v2 Announce Type: replace Abstract: Immersive affective content generation aims to create visually compelling VR imagery with controllable emotional nuance, yet existing methods typica

End-to-End Learning of Safe Optimal Feedback Control in High Dimensions with Control Barrier Function Layers

SafetyDGX agent

arXiv:2607.20674v1 Announce Type: new Abstract: We consider the problem of learning high-dimensional semi-global feedback controllers under hard safety constraints enforced by control barrier function

Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs

SafetyDGX agent

arXiv:2607.20814v1 Announce Type: new Abstract: The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning models remains con- strained by limited i

Environment-Aware Channel Inference via Cross-Modal Flow: From Multimodal Sensing to Wireless Channels

SafetyDGX agent

arXiv:2512.04966v2 Announce Type: replace-cross Abstract: Accurate channel state information (CSI) underpins reliable and efficient wireless communication. However, acquiring CSI via pilot estimation

Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog

SafetyDGX agent

arXiv:2607.21412v1 Announce Type: new Abstract: Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning, especially i

Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective

SafetyDGX agent

arXiv:2604.25077v2 Announce Type: replace Abstract: Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on examples

EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL

SafetyDGX agent

arXiv:2607.20489v1 Announce Type: new Abstract: Text-to-SQL has advanced rapidly with large language models, but complex database queries still require reasoning beyond one-shot generation, including

excellent

SafetyDGX agent

excellent For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen

Expert Behavior Prior Reinforcement Learning

SafetyDGX agent

arXiv:2607.21302v1 Announce Type: new Abstract: Behavior prior reinforcement learning (BPRL) has emerged as a promising paradigm to improve sample efficiency in online reinforcement learning (RL) by l

← Previous
1…2829303132…212
Next →