AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlog
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,629 results
Model Releases

GEO-Bench: Benchmarking Ranking Manipulation in Generative Engine Optimization

DGX agent

arXiv:2605.29107v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly rank products, documents, and recommendations for user queries, which makes manipulating these rankings a gr

model-releasesarxiv-cs-ai
29 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

GPIC: A Giant Permissive Image Corpus for Visual Generation

DGX agent

arXiv:2605.30341v1 Announce Type: cross Abstract: Studying scalable methods for visual generative modeling requires large, accessible, and stable datasets. We introduce GPIC, a Giant Permissive Image

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Gram: Assessing sabotage propensities via automated alignment auditing

DGX agent

arXiv:2605.30322v1 Announce Type: cross Abstract: We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models ac

model-releasesarxiv-cs-ai
29 May 2026
Safety

GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German

DGX agent

arXiv:2605.30214v1 Announce Type: new Abstract: Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about referenc

safetyarxiv-cs-cl
29 May 2026
Model Releases

How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever l…

DGX agent

How lucky are you to have been born when and where you are? Had Opus 4.8 in Claude Code whip up a new visualization of all humans who ever lived. In addition to being neat, it is an interesting test o

model-releasesethan-mollick--x
29 May 2026
Agents

Human-in-the-Loop Swarms: A Bionic Swarm Approach to Real-World Soil Mapping

DGX agent

arXiv:2605.29091v1 Announce Type: new Abstract: Swarm and field robotics face significant barriers to real-world validation due to the high cost and development time to deploy hardware. This paper int

agentsarxiv-cs-ro
29 May 2026
Model Releases

Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles

DGX agent

arXiv:2605.29473v1 Announce Type: cross Abstract: Language models are increasingly being deployed for conversational support in informal caregiving contexts, where interactions often extend beyond inf

model-releasesarxiv-cs-ai
29 May 2026
Applications

LsrIF: Enhancing Logic-Structured Instruction Following of Large Language Models

DGX agent

arXiv:2601.06431v3 Announce Type: replace Abstract: Instruction following is critical for large language models, yet real-world instructions often involve multiple constraints with logical structures,

applicationsarxiv-cs-ai
29 May 2026
Model Releases

MAGA-Bench: Machine-Augment-Generated Text via Alignment Detection Benchmark

DGX agent

arXiv:2601.04633v2 Announce Type: replace Abstract: Machine-Generated Text (MGT) is becoming increasingly difficult to distinguish from Human-Written Text (HWT). This trend has exacerbated malicious a

model-releasesarxiv-cs-cl
29 May 2026
Tutorials

MonoDuo: Using One Robot Arm to Learn Bimanual Policies

DGX agent

arXiv:2605.29298v1 Announce Type: new Abstract: Bimanual coordination is essential for many real-world manipulation tasks, yet learning bimanual robot policies is limited by the scarcity of bimanual r

tutorialsarxiv-cs-ro
29 May 2026
Agents

MOOSE-Copilot: A Web-Based Interactive Assistant for Unified Exploratory and Fine-Grained Scientific Hypothesis Discovery

DGX agent

arXiv:2605.29475v1 Announce Type: cross Abstract: Large language models (LLMs) show remarkable potential in scientific hypothesis discovery. However, existing approaches face two critical limitations:

agentsarxiv-cs-ai
29 May 2026
Model Releases

Multimodal LLMs See Sentiment

DGX agent

arXiv:2508.16873v3 Announce Type: replace Abstract: Understanding how visual content conveys sentiment is increasingly important in a digital landscape dominated by imagery. However, sentiment percept

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs

DGX agent

arXiv:2605.29300v1 Announce Type: cross Abstract: Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses ar

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Personalized Turn-Level User Conversation Satisfaction Benchmark

DGX agent

arXiv:2605.29711v1 Announce Type: cross Abstract: User satisfaction with AI assistants is highly personalized: the same response may satisfy one user but disappoint another depending on what each user

model-releasesarxiv-cs-ai
29 May 2026
Safety

Practitioner Beliefs and Behaviors in AI-Enhanced Education: DOT Framework Survey Evidence

DGX agent

arXiv:2605.29041v1 Announce Type: new Abstract: This study reports findings from a cross-sectional survey (n = 72) of higher education practitioners examining beliefs, behaviors, and institutional con

safetyarxiv-cs-ai
29 May 2026
Applications

Prompts (verbatim): assume a universal veil of ignorance and you could be born as any human who has ever lived in history, what are the most…

DGX agent

Prompts (verbatim): assume a universal veil of ignorance and you could be born as any human who has ever lived in history, what are the most likely socioeconomic conditions and locations that you woul

applicationsethan-mollick--x
29 May 2026
Safety

Provably Secure Agent Guardrail

DGX agent

arXiv:2605.29251v1 Announce Type: new Abstract: As large language models transition from bounded generative engines to agents with expansive execution privileges, AI going out of control precipitates

safetyarxiv-cs-ai
29 May 2026
Model Releases

PTCG-Bench: Can LLM Agents Master Pokemon Trading Card Game?

DGX agent

arXiv:2605.29653v1 Announce Type: new Abstract: Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds. Autonomous agents require sim

model-releasesarxiv-cs-ai
29 May 2026
Safety

Quantum-Enhanced Adversarial Robustness in Artificial Intelligence

DGX agent

arXiv:2605.28899v1 Announce Type: cross Abstract: Artificial Intelligence has achieved remarkable success across diverse application domains. However, its vulnerability to adversarial attacks poses si

safetyarxiv-cs-ai
29 May 2026
Model Releases

RAISE: RAG Design as an Architecture Search Problem

DGX agent

arXiv:2605.30029v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems expose numerous design choices spanning query rewriting, chunking, retrieval depth, reranking, and context

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Realistic honeypot evaluations for scheming propensity

DGX agent

arXiv:2605.29729v1 Announce Type: new Abstract: We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming

model-releasesarxiv-cs-lg
29 May 2026
Safety

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games

DGX agent

arXiv:2602.20141v2 Announce Type: replace Abstract: Mean Field Games (MFGs) provide a principled framework for modelling interactions in large population systems. However, algorithmic progress has bee

safetyarxiv-cs-ai
29 May 2026
Safety

Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents

DGX agent

arXiv:2605.28850v1 Announce Type: new Abstract: We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments. Using TradeArena, an

safetyarxiv-cs-lg
29 May 2026
Agents

// Scaling Laws for Agent Harnesses // If you build agent harnesses, this one is worth your time. (bookmark it) Most harness tuning treats e…

DGX agent

// Scaling Laws for Agent Harnesses // If you build agent harnesses, this one is worth your time. (bookmark it) Most harness tuning treats every token and tool call as if volume is all that counts. Ne

agentsdair-ai--x
29 May 2026
Agents

self-modification

DGX agent

Self-modification in AI systems refers to the capability of an artificial intelligence to alter its own code, parameters, or behavior patterns without external intervention. This concept, discussed by

agentsyohei-nakajima--x
29 May 2026
Model Releases

TAE: Target-aware enhancer for nighttime UAV tracking

DGX agent

arXiv:2605.29558v1 Announce Type: new Abstract: Severe image degradation under low-light nighttime conditions constitutes a core bottleneck preventing all-day applications for UAV-based single object

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Temporal Stability and Few-Shot Prompting in Math Task Assessment

DGX agent

arXiv:2605.30151v1 Announce Type: new Abstract: As AI tools become increasingly integrated into educational contexts, questions arise about both their stability over time and their responsiveness to p

model-releasesarxiv-cs-ai
29 May 2026
Tutorials

Tip on Grok Build + Grok 4.3 VLM One of my critical tasks is to keep Grok VLM in the loop. Throwing a default system prompt usually yields p…

DGX agent

Tip on Grok Build + Grok 4.3 VLM One of my critical tasks is to keep Grok VLM in the loop. Throwing a default system prompt usually yields poor results due to lack of context. Here is how to scale: -

tutorialselon-musk--x
29 May 2026
Model Releases

$65B private round More than double the size of the largest IPO ever

DGX agent

65B private round More than double the size of the largest IPO ever We've raised 65 billion in Series H funding at a $965 billion post-money valuation, led by @AltimeterCap, Dragoneer, @Greenoaks, and

model-releasesjeremy-howard--x
28 May 2026
Agents

AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models

DGX agent

arXiv:2605.27873v1 Announce Type: new Abstract: AI models underpin data-centric applications from image and text processing to scientific discovery in biology, physics, and chemistry. Yet developing t

agentsarxiv-cs-ai
28 May 2026
Tools

[AINews] Cognition raises 1B in 26B Series D

DGX agent

Cognition, an AI company, raised 1 billion in Series D funding at a 26 billion valuation, indicating significant investor confidence in its technology and business model. The funding round reflects th

toolslatent-space
28 May 2026
Model Releases

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation

DGX agent

arXiv:2605.28642v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment par

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Benchmarking Fairness in Spiking Neural Networks: Data Bias, Spurious Features, and Hardware Effects

DGX agent

arXiv:2605.27407v1 Announce Type: cross Abstract: Evaluating fairness in Spiking Neural Networks (SNNs) demands rigorous benchmarks that reflect real-world complexities, yet existing assessments remai

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week @every and our verdict is they could've ju…

DGX agent

BREAKING: Anthropic just dropped Opus 4.8—and it is a MONSTER We've been testing for about a week @every and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check:

model-releasesthariq--x
28 May 2026
Model Releases

Category-Level 3D Correspondence in Camera Space via Morphable Object Priors

DGX agent

arXiv:2605.28257v1 Announce Type: new Abstract: Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category-level pose estim

model-releasesarxiv-cs-cv
28 May 2026
Safety

Causal Direct Preference Optimization for Distributionally Robust Generative Recommendation

DGX agent

arXiv:2603.22335v2 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) guides large language models (LLMs) to generate recommendations aligned with user historical behavior dis

safetyarxiv-cs-ai
28 May 2026
Tutorials

Checking Fact with Better Retrieval: Dynamic Contrastive Learning for Evidence Retrieval

DGX agent

arXiv:2605.27449v1 Announce Type: cross Abstract: In the field of multimodal fact checking, the accuracy of retrieving evidence from different modalities has a significant impact on the downstream cla

tutorialsarxiv-cs-ai
28 May 2026
Model Releases

Claude Opus 4.8: 'a modest but tangible improvement'

DGX agent

Anthropic shipped Claude Opus 4.8 today. My favourite thing about it is this note in the release announcement: Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor. Ther

model-releasessimon-willison
28 May 2026
Model Releases

Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models

DGX agent

arXiv:2512.00349v2 Announce Type: replace Abstract: Are frontier AI systems becoming more capable? Certainly. Yet such progress is not an unalloyed blessing but rather a Trojan horse: behind their per

model-releasesarxiv-cs-ai
28 May 2026
Agents

Defending LLM-based Multi-Agent Systems Against Cooperative Attacks with Sentence-Level Rectification

DGX agent

arXiv:2605.28104v1 Announce Type: new Abstract: Recent years have witnessed the rapid development of Large Language Model-based Multi-Agent Systems (MAS), which excel at collaborative decision-making

agentsarxiv-cs-ai
28 May 2026
Agents

DeltaMCP: Incremental Regeneration via Spec-Aware Transformation for MCP servers

DGX agent

arXiv:2605.28148v1 Announce Type: cross Abstract: The rapid development of LLMs coupled with the introduction of Model Context Protocol (MCP) has revolutionized how intelligent agents interact with AP

agentsarxiv-cs-ai
28 May 2026
Safety

Did AI just solve math? Cal Newport podcast: https://youtu.be/fhZRWZ6J4k4?si=Gzut_5x856n5DW6o

DGX agent

Cal Newport discusses whether recent AI advances represent a breakthrough in mathematical problem-solving capabilities, examining the implications of AI systems' improved performance on complex mathem

safetygary-marcus--x
28 May 2026
Model Releases

Do Agents Think Deeper? A Mechanistic Investigation of Layer-Wise Dynamics in Sequential Planning

DGX agent

arXiv:2605.27935v1 Announce Type: new Abstract: Recent mechanistic studies suggest that large language models (LLMs) may utilize their depth inefficiently in standard single-turn tasks. Whether this s

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Federated Learning for Multivariate Time Series Anomaly Detection in Industrial Automation

DGX agent

arXiv:2605.27486v1 Announce Type: new Abstract: Federated learning (FL) has broadened the horizon for multivariate time series anomaly detection (MTSAD). However, benchmarking such anomaly detection m

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Finding Miscompiles for Fun, Not Profit Or: You don’t need access to Claude Mythos to spend $10,000 in an afternoon https://newsletter.semia…

DGX agent

This post discusses how compiler bugs or miscompiles can be discovered and exploited without expensive AI systems, illustrating that significant computational costs can be incurred quickly when identi

model-releasesdylan-patel--x
28 May 2026
Safety

fun thread on consciousness with @Grimezsz:

DGX agent

Gary Marcus shares a discussion thread about consciousness with user @Grimezsz on X (formerly Twitter), likely exploring philosophical, scientific, or technical perspectives on consciousness and relat

safetygary-marcus--x
28 May 2026
Model Releases

Graph Neural Networks for Source Detection: A Review and Benchmark Study

DGX agent

arXiv:2512.20657v2 Announce Type: replace-cross Abstract: The source detection problem arises when an epidemic process unfolds over a contact network, and the objective is to identify its point of ori

model-releasesarxiv-cs-lg
28 May 2026
Agents

Heterogeneous Multi-Agent Modeling for Measurement and Network Analysis of the Data Service Market

DGX agent

arXiv:2605.27433v1 Announce Type: cross Abstract: With the increasing complexity of collaboration among various social entities and user demands, the factors affecting the stable development of the da

agentsarxiv-cs-ai
28 May 2026
← Previous
1…496497498499500…534
Next →