AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “agents”

GridTimelineEvolution
17,599 results
17 Apr 2026

Can Large Language Models Detect Methodological Flaws? Evidence from Gesture Recognition for UAV-Based Rescue Operation Based on Deep Learning

ApplicationsDGX agent

arXiv:2604.14161v1 Announce Type: new Abstract: Reliable evaluation is essential in machine learning research, yet methodological flaws-particularly data leakage-continue to undermine the validity of

Chain of Modality: From Static Fusion to Dynamic Orchestration in Omni-MLLMs

SafetyDGX agent

arXiv:2604.14520v1 Announce Type: new Abstract: Omni-modal Large Language Models (Omni-MLLMs) promise a unified integration of diverse sensory streams. However, recent evaluations reveal a critical pe

Cognitive Alpha Mining via LLM-Driven Code-Based Evolution

ResearchDGX agent

arXiv:2511.18850v2 Announce Type: replace Abstract: Discovering effective predictive signals, or 'alphas,' from financial data with high dimensionality and extremely low signal-to-noise ratio remains

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Contract-Coding: Towards Repo-Level Generation via Structured Symbolic Paradigm

Model ReleasesDGX agent

arXiv:2604.13100v1 Announce Type: cross Abstract: The shift toward intent-driven software engineering (often termed 'Vibe Coding') exposes a critical Context-Fidelity Trade-off: vague user intents ove

Generative Augmented Inference

ResearchDGX agent

arXiv:2604.14575v1 Announce Type: new Abstract: Data-driven operations management often relies on parameters estimated from costly human-generated labels. Recent advances in large language models (LLM

HRDexDB: A Large-Scale Dataset of Dexterous Human and Robotic Hand Grasps

Model ReleasesDGX agent

arXiv:2604.14944v1 Announce Type: cross Abstract: We present HRDexDB, a large-scale, multi-modal dataset of high-fidelity dexterous grasping sequences featuring both human and diverse robotic hands. U

in retrospect putting the slop cannons (@_lopopolo) on @aiDotEngineer talks day 1 and putting the grown ups (@badlogicgames) on talks day 2 …

ToolsDGX agent

in retrospect putting the slop cannons (@_lopopolo) on @aiDotEngineer talks day 1 and putting the grown ups (@badlogicgames) on talks day 2 is working out pretty well for faithfully representing the m

In @steipete's latest State of the Claw, he gives an update on 5 months of @OpenClaw and some behind the scenes on what it's like maintainin…

ToolsDGX agent

In @steipete's latest State of the Claw, he gives an update on 5 months of @OpenClaw and some behind the scenes on what it's like maintaining the fastest growing open source of all time: https://www.y

MARCA: A Checklist-Based Benchmark for Multilingual Web Search

Model ReleasesDGX agent

arXiv:2604.14448v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select rel

Mechanistic Decoding of Cognitive Constructs in LLMs

Model ReleasesDGX agent

arXiv:2604.14593v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate increasingly sophisticated affective capabilities, the internal mechanisms by which they process complex

Model-Based Reinforcement Learning under Random Observation Delays

ApplicationsDGX agent

arXiv:2509.20869v2 Announce Type: replace Abstract: Delays frequently occur in real-world environments, yet standard reinforcement learning (RL) algorithms often assume instantaneous perception of the

Oh look! Anthropic's entire 'we are delaying Mythos' narrative was marketing hogwash. Kudos to FT for confirming what was obvious. Anthropic…

HardwareDGX agent

Oh look! Anthropic's entire 'we are delaying Mythos' narrative was marketing hogwash. Kudos to FT for confirming what was obvious. Anthropic simply doesn't have the compute. FT: 'Multiple people with

OmniCompliance-100K: A Multi-Domain, Rule-Grounded, Real-World Safety Compliance Dataset

SafetyDGX agent

arXiv:2603.13933v2 Announce Type: replace Abstract: Ensuring the safety and compliance of large language models (LLMs) is of paramount importance. However, existing LLM safety datasets often rely on a

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies

Model ReleasesDGX agent

arXiv:2604.15151v1 Announce Type: new Abstract: Large language models have demonstrated strong performance on general-purpose programming tasks, yet their ability to generate executable algorithmic tr

RECOVER: Designing a Large Language Model-based Remote Patient Monitoring System for Postoperative Gastrointestinal Cancer Care

SafetyDGX agent

arXiv:2502.05740v2 Announce Type: replace-cross Abstract: Cancer surgery is a key treatment for gastrointestinal (GI) cancers, a group of cancers that account for more than 35% of cancer-related death

SPAGBias: Uncovering and Tracing Structured Spatial Gender Bias in Large Language Models

SafetyDGX agent

arXiv:2604.14672v1 Announce Type: new Abstract: Large language models (LLMs) are being increasingly used in urban planning, but since gendered space theory highlights how gender hierarchies are embedd

The Autocorrelation Blind Spot: Why 42% of Turn-Level Findings in LLM Conversation Analysis May Be Spurious

SafetyDGX agent

arXiv:2604.14414v1 Announce Type: new Abstract: Turn-level metrics are widely used to evaluate properties of multi-turn human-LLM conversations, from safety and sycophancy to dialogue quality. However

The Inference Cloud Memory Layer: A Technical Dive into DigitalOcean Managed Databases

IndustryDGX agent

DigitalOcean's Inference Cloud Memory Layer is a technical architecture component designed to optimize database performance by implementing an in-memory caching layer for faster data access and reduce

The PICCO Framework for Large Language Model Prompting: A Taxonomy and Reference Architecture for Prompt Structure

SafetyDGX agent

arXiv:2604.14197v1 Announce Type: new Abstract: Large language model (LLM) performance depends heavily on prompt design, yet prompt construction is often described and applied inconsistently. Our purp

VoxSafeBench: Not Just What Is Said, but Who, How, and Where

SafetyDGX agent

arXiv:2604.14548v1 Announce Type: cross Abstract: As speech language models (SLMs) transition from personal devices into shared, multi-user environments, their responses must account for far more than

16 Apr 2026

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models

ApplicationsDGX agent

arXiv:2511.10946v3 Announce Type: replace Abstract: Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world

Activation-Guided Local Editing for Jailbreaking Attacks

SafetyDGX agent

arXiv:2508.00555v2 Announce Type: replace-cross Abstract: Jailbreaking is an essential adversarial technique for red-teaming these models to uncover and patch security flaws. However, existing jailbre

Alignment as Institutional Design: From Behavioral Correction to Transaction Structure in Intelligent Systems

SafetyDGX agent

arXiv:2604.13079v1 Announce Type: cross Abstract: Current AI alignment paradigms rely on behavioral correction: external supervisors (e.g., RLHF) observe outputs, judge against preferences, and adjust

also available on the Claude Blog: https://claude.com/blog/using-claude-code-session-management-and-1m-context

Model ReleasesDGX agent

Claude Code's session management capabilities and 1 million token context window are highlighted in this post, which references an official Anthropic blog entry. The feature allows developers to maint

Automated co-design of high-performance thermodynamic cycles via graph-based hierarchical reinforcement learning

ResearchDGX agent

arXiv:2604.13133v1 Announce Type: new Abstract: Thermodynamic cycles are pivotal in determining the efficacy of energy conversion systems. Traditional design methodologies, which rely on expert knowle

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning

SafetyDGX agent

arXiv:2604.13804v1 Announce Type: new Abstract: The rapid evolution of multimodal large models has revolutionized the simulation of diverse characters in speech dialogue systems, enabling a novel inte

Databricks on Google Cloud: Innovate Faster. Smarter. Together.

Model ReleasesDGX agent

Databricks and Google Cloud have partnered to enable organizations to build and deploy data and AI solutions more efficiently. The collaboration integrates Databricks' lakehouse platform with Google C

Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation

ResearchDGX agent

arXiv:2604.13088v1 Announce Type: new Abstract: In sparse termination rewards, intra-group comparisons have become the dominant paradigm for fine-tuning reasoning models via reinforcement learning. Ho

Designing synthetic datasets for the real world: Mechanism design and reasoning from first principles

ApplicationsDGX agent

This Google Research work presents guidelines for synthetic data mechanism design and provides insights into generating and evaluating synthetic data at scale. The research introduces a reasoning-driv

Developer tooling startup Expo nabs $45M investment

IndustryDGX agent

Expo, the developer of a popular open-source tool for building cross-platform applications, today announced that it has raised 45 million in funding. Developers often implement web application interfa

Dual-Enhancement Product Bundling: Bridging Interactive Graph and Large Language Model

ResearchDGX agent

arXiv:2604.14030v1 Announce Type: new Abstract: Product bundling boosts e-commerce revenue by recommending complementary item combinations. However, existing methods face two critical challenges: (1)

EmbodiedClaw: Conversational Workflow Execution for Embodied AI Development

Model ReleasesDGX agent

arXiv:2604.13800v1 Announce Type: new Abstract: Embodied AI research is increasingly moving beyond single-task, single-environment policy learning toward multi-task, multi-scene, and multi-model setti

ESCAPE: Episodic Spatial Memory and Adaptive Execution Policy for Long-Horizon Mobile Manipulation

Model ReleasesDGX agent

arXiv:2604.13633v1 Announce Type: new Abstract: Coordinating navigation and manipulation with robust performance is essential for embodied AI in complex indoor environments. However, as tasks extend o

I edited the intro because I realized I buried the lede originally- The 1M context window is a double-edged sword. It allows Claude to do mo…

Model ReleasesDGX agent

I edited the intro because I realized I buried the lede originally- The 1M context window is a double-edged sword. It allows Claude to do more complex tasks but it can also leads to more context pollu

I love to explore solutions with it so I often ask it to brainstorm with me, then choose an option and rewind to implement it. When I'm read…

ToolsDGX agent

I love to explore solutions with it so I often ask it to brainstorm with me, then choose an option and rewind to implement it. When I'm ready, I ask it to interview me to figure out what’s in my head

IndicDB -- Benchmarking Multilingual Text-to-SQL Capabilities in Indian Languages

Model ReleasesDGX agent

arXiv:2604.13686v1 Announce Type: new Abstract: While Large Language Models (LLMs) have significantly advanced Text-to-SQL performance, existing benchmarks predominantly focus on Western contexts and

LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models

Model ReleasesDGX agent

arXiv:2511.11334v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has not been matched by their evaluation in low-resource languages, especially Southeast Asian

Learn more about our journey https://youtu.be/aaBRSWWB_tI

TutorialsDGX agent

Replit shared a YouTube video detailing their company's journey and history of development. The video likely covers key milestones, founding story, product evolution, and the team's vision for their c

Lossless Prompt Compression via Dictionary-Encoding and In-Context Learning: Enabling Cost-Effective LLM Analysis of Repetitive Data

Model ReleasesDGX agent

arXiv:2604.13066v1 Announce Type: new Abstract: In-context learning has established itself as an important learning paradigm for Large Language Models (LLMs). In this paper, we demonstrate that LLMs c

Multi-Dimensional Knowledge Profiling with Large-Scale Literature Database and Hierarchical Retrieval

SafetyDGX agent

arXiv:2601.15170v2 Announce Type: replace Abstract: The rapid expansion of research across machine learning, vision, and language has produced a volume of publications that is increasingly difficult t

Olfactory pursuit: catching a moving odor source in complex flows

Model ReleasesDGX agent

arXiv:2604.13121v1 Announce Type: new Abstract: Locating and intercepting a moving target from possibly delayed, intermittent sensory signals is a paradigmatic problem in decision-making under uncerta

🎬 Ollama Gemma Day Recap: SGLang at the Ollama Gemma 4 Party in Palo Alto 🍾 Last night, @ollama hosted a packed Gemma Day at the Palo Alto…

Model ReleasesDGX agent

🎬 Ollama Gemma Day Recap: SGLang at the Ollama Gemma 4 Party in Palo Alto 🍾 Last night, @ollama hosted a packed Gemma Day at the Palo Alto office alongside the @GoogleDeepMind Gemma team. SGLang was i

Peer-Predictive Self-Training for Language Model Reasoning

Model ReleasesDGX agent

arXiv:2604.13356v1 Announce Type: new Abstract: Mechanisms for continued self-improvement of language models without external supervision remain an open challenge. We propose Peer-Predictive Self-Trai

qwen3.6 is out

Local AiDGX agent

Qwen 3.6 Plus Preview is Alibaba's next-generation large language model released on March 30-31, 2026 , and the first open-weight variant was released following the February Qwen 3.5 series, prioritiz

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Model ReleasesDGX agent

arXiv:2604.13602v1 Announce Type: new Abstract: Reinforcement Learning from Human Feedback (RLHF) and related alignment paradigms have become central to steering large language models (LLMs) and multi

RPS: Information Elicitation with Reinforcement Prompt Selection

Model ReleasesDGX agent

arXiv:2604.13817v1 Announce Type: new Abstract: Large language models (LLMs) have shown remarkable capabilities in dialogue generation and reasoning, yet their effectiveness in eliciting user-known bu

Seek-and-Solve: Benchmarking MLLMs for Visual Clue-Driven Reasoning in Daily Scenarios

Model ReleasesDGX agent

arXiv:2604.14041v1 Announce Type: new Abstract: Daily scenarios are characterized by visual richness, requiring Multimodal Large Language Models (MLLMs) to filter noise and identify decisive visual cl

Self-adaptive Multi-Access Edge Architectures: A Robotics Case

SafetyDGX agent

arXiv:2604.13542v1 Announce Type: new Abstract: The growth of compute-intensive AI tasks highlights the need to mitigate the processing costs and improve performance and energy efficiency. This necess

Shocking result on my pelican benchmark this morning, I got a better pelican from a 21GB local Qwen3.6-35B-A3B running on my laptop than I d…

Model ReleasesDGX agent

Shocking result on my pelican benchmark this morning, I got a better pelican from a 21GB local Qwen3.6-35B-A3B running on my laptop than I did from the new Opus 4.7! Qwen on the left, Opus on the righ

Sorry for the long wait, everyone! As I said, Qwen is going to keep open-sourcing!

Model ReleasesDGX agent

Sorry for the long wait, everyone! As I said, Qwen is going to keep open-sourcing! ⚡ Meet Qwen3.6-35B-A3B:Now Open-Source!🚀🚀 A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. 🔥 Agen

this makes sense all else held equal, concentrated fund = less risky bets we've found that 36 companies is the max number of companies we ca…

TutorialsDGX agent

this makes sense all else held equal, concentrated fund = less risky bets we've found that 36 companies is the max number of companies we can fit into a fund while claiming that a single company at 1B

TIP: Token Importance in On-Policy Distillation

Model ReleasesDGX agent

arXiv:2604.14084v1 Announce Type: new Abstract: On-policy knowledge distillation (OPD) trains a student on its own rollouts under token-level supervision from a teacher. Not all token positions matter

With 4.7 you can push a lot further with one prompt. That means multi-file changes, ambiguous debugging, code review across a whole service.…

ToolsDGX agent

With 4.7 you can push a lot further with one prompt. That means multi-file changes, ambiguous debugging, code review across a whole service. The stuff you used to break into small chunks because the m

15 Apr 2026

A Sanity Check on Composed Image Retrieval

Model ReleasesDGX agent

arXiv:2604.12904v1 Announce Type: new Abstract: Composed Image Retrieval (CIR) aims to retrieve a target image based on a query composed of a reference image, and a relative caption that specifies the

Accelerating decode-heavy LLM inference with speculative decoding on AWS Trainium and vLLM

TutorialsDGX agent

Speculative decoding is a technique used to accelerate the slow, sequential token generation (decode stage) in LLM inference. This method significantly reduces latency and improves hardware utilizatio

Analysis: nearly 90 schools and 600 students globally have been impacted by AI-generated deepfake nudes; North America had nearly 30 reported cases since 2023 (Matt Burgess/Wired)

IndustryDGX agent

Matt Burgess / Wired: Analysis: nearly 90 schools and 600 students globally have been impacted by AI-generated deepfake nudes; North America had nearly 30 reported cases since 2023 — An analysis by WI

Been waiting a month for Anthropic to answer a simple usage question about Claude Code subscriptions Have I been ghosted

Model ReleasesDGX agent

Been waiting a month for Anthropic to answer a simple usage question about Claude Code subscriptions Have I been ghosted Can I get some questions answered by someone at Anthropic? 1. Can you use an OA

Boston Dynamics’ robot dog now reads gauges and thermometers with Google's AI

IndustryDGX agent

Boston Dynamics has integrated Google's Gemini and Gemini Robotics-ER 1.6 into its Orbit software platform, specifically its AI Visual Inspection systems, which analyze images captured by the Spot rob

Cal.com, which provides scheduling software, is moving its core open-source codebase to a closed repository, citing the dangers of AI hacking its open code (Steven Vaughan-Nichols/ZDNET)

IndustryDGX agent

Steven Vaughan-Nichols / ZDNET: Cal.com, which provides scheduling software, is moving its core open-source codebase to a closed repository, citing the dangers of AI hacking its open code — ZDNET's ke

Deep QP Safety Filter: Model-free Learning for Reachability-based Safety Filter

SafetyDGX agent

arXiv:2601.21297v2 Announce Type: replace Abstract: We introduce Deep QP Safety Filter, a fully data-driven safety layer for black-box dynamical systems. Our method learns a Quadratic-Program (QP) saf

← Previous
1…287288289290291…294
Next →