AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,599 results
23 Apr 2026

Things have been degrading super fast in Claude Code. I still use Claude Code, but my default is now Codex. I still prefer Opus models for c…

Model ReleasesDGX agent

Things have been degrading super fast in Claude Code. I still use Claude Code, but my default is now Codex. I still prefer Opus models for coding, and so I will try again with the fixes. I appreciate

US Commerce Secretary Howard Lutnick says Nvidia has yet to sell H200 chips to Chinese companies and that the Chinese government has not approved such purchases (Alexandra Alper/Reuters)

HardwareDGX agent

Alexandra Alper / Reuters: US Commerce Secretary Howard Lutnick says Nvidia has yet to sell H200 chips to Chinese companies and that the Chinese government has not approved such purchases — Nvidia's (

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2604.20398v1 Announce Type: new Abstract: While Large Language Models (LLMs) excel at function-level code generation, project-level tasks such as generating functional and visually aesthetic mul

We’re resetting usage limits for subscribers. Thank you so much for your feedback and patience!

ToolsDGX agent

Boris Cherny announced that usage limits are being reset for subscribers in response to user feedback and patience. The post suggests this is a policy adjustment made by a service or platform to addre

22 Apr 2026

Accelerating Optimization and Machine Learning through Decentralization

TutorialsDGX agent

arXiv:2604.19518v1 Announce Type: new Abstract: Decentralized optimization enables multiple devices to learn a global machine learning model while each individual device only has access to its local d

Announcing Spanner Omni: Your infrastructure, Google’s innovation

Model ReleasesDGX agent

Today, we announced the preview of Spanner Omni, a downloadable version of Spanner, that expands its industry-leading distributed database capabilities beyond Google Cloud. This enables enterprises to

ASVSim (AirSim for Surface Vehicles): A High-Fidelity Simulation Framework for Autonomous Surface Vehicle Research

SafetyDGX agent

arXiv:2506.22174v2 Announce Type: replace-cross Abstract: The transport industry has recently shown significant interest in unmanned surface vehicles (USVs), specifically for port and inland waterway

BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental Design

ResearchDGX agent

arXiv:2508.21184v3 Announce Type: replace-cross Abstract: We propose a general-purpose approach for improving the ability of large language models (LLMs) to intelligently and adaptively gather informa

Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications

SafetyDGX agent

arXiv:2604.19281v1 Announce Type: cross Abstract: The use of Large Language Models (LLMs) to support patients in addressing medical questions is becoming increasingly prevalent. However, most of the m

Do Emotions Influence Moral Judgment in Large Language Models?

SafetyDGX agent

arXiv:2604.19125v1 Announce Type: new Abstract: Large language models have been extensively studied for emotion recognition and moral reasoning as distinct capabilities, yet the extent to which emotio

Evaluation-driven Scaling for Scientific Discovery

Local AiDGX agent

arXiv:2604.19341v1 Announce Type: cross Abstract: Language models are increasingly used in scientific discovery to generate hypotheses, propose candidate solutions, implement systems, and iteratively

EVPO: Explained Variance Policy Optimization for Adaptive Critic Utilization in LLM Post-Training

SafetyDGX agent

arXiv:2604.19485v1 Announce Type: cross Abstract: Reinforcement learning (RL) for LLM post-training faces a fundamental design choice: whether to use a learned critic as a baseline for policy optimiza

GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes

ResearchDGX agent

arXiv:2604.19411v1 Announce Type: cross Abstract: Understanding road scenes in a geometrically consistent, scene-centric representation is crucial for planning and mapping. We present GOLD-BEV, a fram

LASER: Learning Active Sensing for Continuum Field Reconstruction

SafetyDGX agent

arXiv:2604.19355v1 Announce Type: cross Abstract: High-fidelity measurements of continuum physical fields are essential for scientific discovery and engineering design but remain challenging under spa

Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic

SafetyDGX agent

arXiv:2604.19567v1 Announce Type: new Abstract: Reinforcement learning (RL) as post-training is crucial for enhancing the reasoning ability of large language models (LLMs) in coding and math. However,

Multi-Task Reinforcement Learning for Enhanced Multimodal LLM-as-a-Judge

SafetyDGX agent

arXiv:2603.11665v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have been widely adopted as MLLM-as-a-Judges due to their strong alignment with human judgment across vario

Neuromorphic Continual Learning for Sequential Deployment of Nuclear Plant Monitoring Systems

SafetyDGX agent

arXiv:2604.18611v1 Announce Type: cross Abstract: Anomaly detection in nuclear industrial control systems (ICS) requires continuous, energy-efficient monitoring across multiple subsystems that are oft

OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models

ResearchDGX agent

arXiv:2502.16161v2 Announce Type: replace-cross Abstract: Visually-situated text parsing (VsTP) has recently seen notable advancements, driven by the growing demand for automated document understandin

Paparazzo: Active Mapping of Moving 3D Objects

Model ReleasesDGX agent

arXiv:2604.19556v1 Announce Type: new Abstract: Current 3D mapping pipelines generally assume static environments, which limits their ability to accurately capture and reconstruct moving objects. To a

Reinforcement Learning Enabled Adaptive Multi-Task Control for Bipedal Soccer Robots

ResearchDGX agent

arXiv:2604.19104v1 Announce Type: cross Abstract: Developing bipedal football robots in dynamiccombat environments presents challenges related to motionstability and deep coupling of multiple tasks, a

RL-ABC: Reinforcement Learning for Accelerator Beamline Control

SafetyDGX agent

arXiv:2604.19146v1 Announce Type: new Abstract: Particle accelerator beamline optimization is a high-dimensional control problem traditionally requiring significant expert intervention. We present RLA

RoboWM-Bench: A Benchmark for Evaluating World Models in Robotic Manipulation

Model ReleasesDGX agent

arXiv:2604.19092v1 Announce Type: cross Abstract: Recent advances in large-scale video world models have enabled increasingly realistic future prediction, raising the prospect of leveraging imagined v

Semantic Interaction Information mediates compositional generalization in latent space

ResearchDGX agent

arXiv:2603.27134v3 Announce Type: replace Abstract: Are there still barriers to generalization once all relevant variables are known? We address this question via a framework that casts compositional

SpaceX partners with Cursor on AI training, floats potential $60B acquisition

IndustryDGX agent

SpaceX Corp. will help Cursor, a venture-backed vibe coding startup, train artificial intelligence models optimized for programming tasks. The companies announced the initiative on Tuesday. According

STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming

SafetyDGX agent

arXiv:2604.18976v1 Announce Type: new Abstract: While Large Language Models (LLMs) are widely used, they remain susceptible to jailbreak prompts that can elicit harmful or inappropriate responses. Thi

This is a really big deal - it's easy to run into nasty bills with Cloud Run if your site attracts aggressive scrapers, spending caps make i…

ToolsDGX agent

This is a really big deal - it's easy to run into nasty bills with Cloud Run if your site attracts aggressive scrapers, spending caps make it a whole lot safer to run small projects on Anouncing Spend

TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process Rewards

SafetyDGX agent

arXiv:2512.07761v3 Announce Type: replace Abstract: Large language models have seen widespread adoption, yet they remain vulnerable to multi-turn jailbreak attacks, threatening their safe deployment.

21 Apr 2026

Amortized Inverse Kinematics via Graph Attention for Real-Time Human Avatar Animation

ResearchDGX agent

arXiv:2604.16629v1 Announce Type: new Abstract: Inverse kinematics (IK) is a core operation in animation, robotics, and biomechanics: given Cartesian constraints, recover joint rotations under a known

Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods

Model ReleasesDGX agent

arXiv:2510.07143v3 Announce Type: replace Abstract: Recent efforts to accelerate inference in Multimodal Large Language Models (MLLMs) have largely focused on visual token compression. The effectivene

Asset Harvester: Extracting 3D Assets from Autonomous Driving Logs for Simulation

SafetyDGX agent

arXiv:2604.18468v1 Announce Type: new Abstract: Closed-loop simulation is a core component of autonomous vehicle (AV) development, enabling scalable testing, training, and safety validation before rea

Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization

SafetyDGX agent

arXiv:2604.17188v1 Announce Type: new Abstract: Multi-role dialogue summarization requires modeling complex interactions among multiple speakers while preserving role-specific information and factual

Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation

ResearchDGX agent

arXiv:2604.17020v1 Announce Type: new Abstract: Static benchmarks for harmful content detection face limitations in scalability and diversity, and may also be affected by contamination from web-scale

Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling

Model ReleasesDGX agent

arXiv:2604.17794v1 Announce Type: new Abstract: The democratization of ubiquitous AI hinges on deploying sophisticated reasoning capabilities on resource-constrained devices. However, Small Language M

CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in China

SafetyDGX agent

arXiv:2510.08986v2 Announce Type: replace Abstract: We introduce CAPC-CG, the Chinese Adaptive Policy Communication (Central Government) Corpus, the first open dataset of Chinese policy directives ann

Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

Model ReleasesDGX agent

arXiv:2507.20409v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting helps models think step by step. But naive CoT breaks down in visually grounded social tasks, where models must per

Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts

SafetyDGX agent

arXiv:2604.18091v1 Announce Type: new Abstract: Recent multimodal large language models have shown promising ability in generating humorous captions for images, yet they still lack stable control over

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies

ResearchDGX agent

arXiv:2503.14324v3 Announce Type: replace-cross Abstract: The differing representation spaces required for visual understanding and generation pose a challenge in unifying them within the autoregressi

Dynamic Emotion and Personality Profiling for Multimodal Deception Detection

SafetyDGX agent

arXiv:2604.17037v1 Announce Type: new Abstract: Deception detection is of great significance for ensuring information security and conducting public opinion analysis, with personality factors and emot

Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation

ResearchDGX agent

arXiv:2502.13637v2 Announce Type: replace Abstract: Human affordance learning investigates contextually relevant novel pose prediction such that the estimated pose represents a valid human action with

Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence

SafetyDGX agent

arXiv:2601.11886v2 Announce Type: replace Abstract: In high-stakes domains like medicine, it may be generally desirable for models to faithfully adhere to the context provided. But what happens if the

Human-Centered Supervision for Sentiment Analysis in Telugu: A Systematic Inquiry Beyond Accuracy

SafetyDGX agent

arXiv:2508.01486v3 Announce Type: replace Abstract: Sentiment analysis for low-resource languages remains challenging in an era where interpretability, human alignment, and fairness are increasingly n

I built a web analytics system to track traffic and events across all my @Replit projects in one place. Drop-in @posthog stack: → Full analy…

ApplicationsDGX agent

I built a web analytics system to track traffic and events across all my @Replit projects in one place. Drop-in @posthog stack: → Full analytics dashboard → Reverse proxy (beat ad blockers) → AI skill

I find that open weights models over-perform on benchmarks compared to actual real-world usage, and Kimi feels like no exception. For exampl…

Model ReleasesDGX agent

I find that open weights models over-perform on benchmarks compared to actual real-world usage, and Kimi feels like no exception. For example, a small amount of use will show that Kimi is not as good

Jupiter-N Technical Report

Model ReleasesDGX agent

arXiv:2604.17429v1 Announce Type: new Abstract: We present Jupiter-N, a hybrid reasoning model post-trained from Nemotron 3 Super, a fully open-source 120 billion parameter LLM. We target three object

K2.6 + hermes = 4 hr session setting up qwen 3.6 training regime on dgx spark with the current autnomous session currently lasting 70+ min w…

Model ReleasesDGX agent

K2.6 + hermes = 4 hr session setting up qwen 3.6 training regime on dgx spark with the current autnomous session currently lasting 70+ min without any prompting. Kimi with hermes is next level. @NousR

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

Model ReleasesDGX agent

arXiv:2603.16120v2 Announce Type: replace Abstract: Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queri

Learning-Based Sparsification of Dynamic Graphs in Robotic Exploration Algorithms

SafetyDGX agent

arXiv:2604.16509v1 Announce Type: cross Abstract: Many robotic exploration algorithms rely on graph structures for frontier-based exploration and dynamic path planning. However, these graphs grow rapi

MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models

SafetyDGX agent

arXiv:2604.17730v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly explored as scalable tools for mental health counseling, yet evaluating their safety remains challenging d

Modeling User Exploration Saturation: When Recommender Systems Should Stop Pushing Novelty

SafetyDGX agent

arXiv:2604.16419v1 Announce Type: cross Abstract: Fairness-aware recommender systems often mitigate bias by increasing exposure to under-represented or long-tail content, commonly through mechanisms t

Neuro-Symbolic Resolution of Recommendation Conflicts in Multimorbidity Clinical Guidelines

Model ReleasesDGX agent

arXiv:2604.17340v1 Announce Type: new Abstract: Clinical guidelines, typically developed by independent specialty societies, inherently exhibit substantial fragmentation, redundancy, and logical contr

NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions

Model ReleasesDGX agent

arXiv:2604.16493v1 Announce Type: cross Abstract: Natural Language to SQL (NL2SQL) technology empowers non-expert users to query relational databases without requiring SQL expertise. While large langu

OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation

Model ReleasesDGX agent

arXiv:2506.05606v5 Announce Type: replace Abstract: Can large language models (LLMs) accurately simulate the next web action of a specific user? While LLMs have shown promising capabilities in generat

Plasticity Loss in Deep Reinforcement Learning: A Survey

SafetyDGX agent

arXiv:2411.04832v3 Announce Type: replace-cross Abstract: Plasticity refers to a network's ability to adapt to changing data distributions, which is crucial for the successful training of deep reinfor

Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?

Model ReleasesDGX agent

arXiv:2604.17338v1 Announce Type: cross Abstract: Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but o

PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues

ResearchDGX agent

arXiv:2604.18354v1 Announce Type: new Abstract: Emotion plays a pivotal role in shaping negotiation outcomes, influencing trust, cooperation, and long-term relationships. Developing negotiation dialog

R3D2: Realistic 3D Asset Insertion via Diffusion for Autonomous Driving Simulation

SafetyDGX agent

arXiv:2506.07826v2 Announce Type: replace Abstract: Validating autonomous driving (AD) systems requires diverse and safety-critical testing, making photorealistic virtual environments essential. Tradi

REALM: Reliable Expertise-Aware Language Model Fine-Tuning from Noisy Annotations

ResearchDGX agent

arXiv:2604.17289v1 Announce Type: new Abstract: Supervised fine-tuning of large language models relies on human-annotated data, yet annotation pipelines routinely involve multiple crowdworkers of hete

Scaling Human-AI Coding Collaboration Requires a Governable Consensus Layer

Model ReleasesDGX agent

arXiv:2604.17883v1 Announce Type: cross Abstract: Vibe coding produces correct, executable code at speed, but leaves no record of the structural commitments, dependencies, or evidence behind it. Revie

ScenarioControl: Vision-Language Controllable Vectorized Latent Scenario Generation

ResearchDGX agent

arXiv:2604.17147v1 Announce Type: new Abstract: We introduce ScenarioControl, the first vision-language control mechanism for learned driving scenario generation. Given a text prompt or an input image

SentiAvatar: Towards Expressive and Interactive Digital Humans

ResearchDGX agent

arXiv:2604.02908v2 Announce Type: replace Abstract: We present SentiAvatar, a framework for building expressive interactive 3D digital humans, and use it to create SuSu, a virtual character that speak

← Previous
1…285286287288289…294
Next →