AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,429 results
21 Apr 2026

VIDS: A Verified Imaging Dataset Standard for Medical AI

Model ReleasesDGX agent

arXiv:2604.17525v1 Announce Type: cross Abstract: Medical imaging AI development is fundamentally dependent on annotated datasets, yet no existing standard provides machine-enforceable validation acro

We are entering an extremely exciting era for open-weight models. Kimi K2.6 now feels like a top agentic model. I took it for a spin via @Fi…

AgentsDGX agent

We are entering an extremely exciting era for open-weight models. Kimi K2.6 now feels like a top agentic model. I took it for a spin via @FireworksAI_HQ fast inference APIs. Kimi K2.6 has impressive a

WeatherArchive-Bench: Benchmarking Retrieval-Augmented Reasoning for Historical Weather Archives

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2510.05336v2 Announce Type: replace Abstract: Historical archives on weather events are collections of enduring primary source records that offer rich, untapped narratives of how societies have

When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators

Model ReleasesDGX agent

arXiv:2602.19946v4 Announce Type: replace Abstract: Recent text-to-image (T2I) diffusion models produce visually stunning images and demonstrate excellent prompt following. But do they perform well as

ZoFia: Zero-Shot Fake News Detection with Entity-Guided Retrieval and Multi-LLM Interaction

SafetyDGX agent

arXiv:2511.01188v2 Announce Type: replace Abstract: The rapid spread of fake news threatens social stability and public trust, highlighting the urgent need for its effective detection. Although large

20 Apr 2026

A Comparative Study on the Impact of Traditional Learning and Interactive Learning on Students' Academic Performance and Emotional Well-Being

TutorialsDGX agent

arXiv:2604.15335v1 Announce Type: cross Abstract: The growing adoption of interactive learning tools in higher education offers new opportunities to enhance student performance and well-being. This st

A methodology to rank importance of frequencies and channels in electromyography data with Decision Tree classifiers

TutorialsDGX agent

arXiv:2604.15353v1 Announce Type: cross Abstract: This study presents a methodology for identifying the most informative frequencies and channels in electromyography (EMG) data to evaluate muscle reco

AI-assisted Protocol Information Extraction For Improved Accuracy and Efficiency in Clinical Trial Workflows

ApplicationsDGX agent

arXiv:2602.00052v2 Announce Type: replace-cross Abstract: Increasing clinical trial protocol complexity, amendments, and challenges around knowledge management create significant burden for trial team

AISysRev -- LLM-based Tool for Title-abstract Screening

Model ReleasesDGX agent

arXiv:2510.06708v3 Announce Type: replace-cross Abstract: Conducting systematic reviews is laborious. In the screening or study selection phase, the number of papers can be overwhelming. Recent resear

Applied Explainability for Large Language Models: A Comparative Study

ApplicationsDGX agent

arXiv:2604.15371v1 Announce Type: cross Abstract: Large language models (LLMs) achieve strong performance across many natural language processing tasks, yet their decision processes remain difficult t

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants

Model ReleasesDGX agent

arXiv:2510.24328v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used to answer everyday questions, yet their performance on culturally grounded and dialectal co

Bilevel Optimization of Agent Skills via Monte Carlo Tree Search

AgentsDGX agent

arXiv:2604.15709v1 Announce Type: new Abstract: Agent exttt{skills} are structured collections of instructions, tools, and supporting resources that help large language model (LLM) agents perform part

Classic study gave 146 economist teams the same dataset & got wildly different answers New paper reruns it with agentic AI. Claude Code & Co…

Model ReleasesDGX agent

Classic study gave 146 economist teams the same dataset & got wildly different answers New paper reruns it with agentic AI. Claude Code & Codex land near the human median, but with far tighter dispers

Cognitive Agency Surrender: Defending Epistemic Sovereignty via Scaffolded AI Friction

SafetyDGX agent

arXiv:2603.21735v2 Announce Type: replace-cross Abstract: The proliferation of Generative Artificial Intelligence has transformed benign cognitive offloading into a systemic risk of cognitive agency s

CTSCAN: Evaluation Leakage in Chest CT Segmentation and a Reproducible Patient-Disjoint Benchmark

Model ReleasesDGX agent

arXiv:2604.15561v1 Announce Type: cross Abstract: Reported chest CT segmentation performance can be strongly inflated when train and test partitions mix slices from the same study. We present CTSCAN,

Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning

ApplicationsDGX agent

arXiv:2604.16029v1 Announce Type: new Abstract: Parallel reasoning enhances Large Reasoning Models (LRMs) but incurs prohibitive costs due to futile paths caused by early errors. To mitigate this, pat

DASB -- Discrete Audio and Speech Benchmark

Model ReleasesDGX agent

arXiv:2406.14294v3 Announce Type: replace-cross Abstract: Discrete audio tokens have recently gained considerable attention for their potential to bridge audio and language processing, enabling multim

DataCenterGym: A Physics-Grounded Simulator for Multi-Objective Data Center Scheduling

Local AiDGX agent

arXiv:2604.15594v1 Announce Type: cross Abstract: Modern datacenters schedule heterogeneous workloads across geo-distributed sites with diverse compute capacities, electricity prices, and thermal cond

DINOv3 Beats Specialized Detectors: A Simple Foundation Model Baseline for Image Forensics

Local AiDGX agent

arXiv:2604.16083v1 Announce Type: new Abstract: With the rapid advancement of deep generative models, realistic fake images have become increasingly accessible, yet existing localization methods rely

Discover and Prove: An Open-source Agentic Framework for Hard Mode Automated Theorem Proving in Lean 4

AgentsDGX agent

arXiv:2604.15839v1 Announce Type: new Abstract: Most ATP benchmarks embed the final answer within the formal statement -- a convention we call 'Easy Mode' -- a design that simplifies the task relative

DyTact: Capturing Dynamic Contacts in Hand-Object Manipulation

SafetyDGX agent

arXiv:2506.03103v2 Announce Type: replace Abstract: Reconstructing dynamic hand-object contacts is essential for realistic manipulation in AI character animation, XR, and robotics, yet it remains chal

Earlier this year Yann LeCun left Meta because Mark Zuckerberg wouldn't bet the company on JEPA. Last week his group dropped the first JEPA …

Model ReleasesDGX agent

Earlier this year Yann LeCun left Meta because Mark Zuckerberg wouldn't bet the company on JEPA. Last week his group dropped the first JEPA that actually trains end-to-end from raw pixels. 15 million

'Excuse me, may I say something...' CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations

Model ReleasesDGX agent

arXiv:2604.15588v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into scientific workflows presents exciting opportunities to accelerate biomedical discovery. However,

Exploitation Over Exploration: Unmasking the Bias in Linear Bandit Recommender Offline Evaluation

SafetyDGX agent

arXiv:2507.18756v2 Announce Type: replace Abstract: Multi-Armed Bandit (MAB) algorithms are widely used in recommender systems that require continuous, incremental learning. A core aspect of MABs is t

Exploring the Capability Boundaries of LLMs in Mastering of Chinese Chouxiang Language

Model ReleasesDGX agent

arXiv:2604.15841v1 Announce Type: new Abstract: While large language models (LLMs) have achieved remarkable success in general language tasks, their performance on Chouxiang Language, a representative

From Articles to Canopies: Knowledge-Driven Pseudo-Labelling for Tree Species Classification using LLM Experts

ApplicationsDGX agent

arXiv:2604.16115v1 Announce Type: new Abstract: Hyperspectral tree species classification is challenging due to limited and imbalanced class labels, spectral mixing (overlapping light signatures from

From Intention to Text: AI-Supported Goal Setting in Academic Writing

SafetyDGX agent

arXiv:2604.15800v1 Announce Type: cross Abstract: This study presents WriteFlow, an AI voice-based writing assistant designed to support reflective academic writing through goal-oriented interaction.

From Zero to Detail: A Progressive Spectral Decoupling Paradigm for UHD Image Restoration with New Benchmark

Model ReleasesDGX agent

arXiv:2604.15654v1 Announce Type: new Abstract: Ultra-high-definition (UHD) image restoration poses unique challenges due to the high spatial resolution, diverse content, and fine-grained structures p

Important to note: that 3x increase for images is entirely due to Opus 4.7 being able to handle higher resolutions. I tried that again with …

ToolsDGX agent

Important to note: that 3x increase for images is entirely due to Opus 4.7 being able to handle higher resolutions. I tried that again with a 682x318 pixel image and it took 314 tokens with Opus 4.7 a

Integrating Graphs, Large Language Models, and Agents: Reasoning and Retrieval

AgentsDGX agent

arXiv:2604.15951v1 Announce Type: new Abstract: Generative AI, particularly Large Language Models, increasingly integrates graph-based representations to enhance reasoning, retrieval, and structured d

Interpretable Traces, Unexpected Outcomes: Investigating the Disconnect in Trace-Based Knowledge Distillation

Model ReleasesDGX agent

arXiv:2505.13792v2 Announce Type: replace-cross Abstract: Recent advances in reasoning-focused Large Language Models (LLMs) have introduced Chain-of-Thought (CoT) traces - intermediate reasoning steps

IPQA: A Benchmark for Core Intent Identification in Personalized Question Answering

Model ReleasesDGX agent

arXiv:2510.23536v2 Announce Type: replace Abstract: Intent identification serves as the foundation for generating appropriate responses in personalized question answering (PQA). However, existing benc

Je suis passé à Découverte de @CBCRadioCanada pour discuter des risques de l’IA, des raisons scientifiques qui expliquent certains des compo…

SafetyDGX agent

Je suis passé à Découverte de @CBCRadioCanada pour discuter des risques de l’IA, des raisons scientifiques qui expliquent certains des comportements inquiétants des modèles de pointe, et des solutions

Kimi has been the most popular model on Fireworks, both out of the box and as a fine-tuning base (including Composer 2) Now Kimi K2.6 is liv…

ToolsDGX agent

Kimi has been the most popular model on Fireworks, both out of the box and as a fine-tuning base (including Composer 2) Now Kimi K2.6 is live with huge jumps (10+%) in coding, long-running agents and

Kimi-K2.6 is on HuggingFace

IndustryDGX agent

Kimi-K2.6 is a language model that has been released on HuggingFace, a popular platform for sharing machine learning models and datasets. The announcement was made by Clem Delangue, likely indicating

Language, Place, and Social Media: Geographic Dialect Alignment in New Zealand

SafetyDGX agent

arXiv:2604.15744v1 Announce Type: new Abstract: This thesis investigates geographic dialect alignment in place-informed social media communities, focussing on New Zealand-related Reddit communities. B

LinuxArena: A Control Setting for AI Agents in Live Production Software Environments

Model ReleasesDGX agent

arXiv:2604.15384v1 Announce Type: cross Abstract: We introduce LinuxArena, a control setting in which agents operate directly on live, multi-service production environments. LinuxArena contains 20 env

llm-openrouter 0.6

ToolsDGX agent

Release: llm-openrouter 0.6 llm openrouter refresh command for refreshing the list of available models without waiting for the cache to expire. I added this feature so I could try Kimi 2.6 on OpenRout

LLMSniffer: Detecting LLM-Generated Code via GraphCodeBERT and Supervised Contrastive Learning

Model ReleasesDGX agent

arXiv:2604.16058v1 Announce Type: cross Abstract: The rapid proliferation of Large Language Models (LLMs) in software development has made distinguishing AI-generated code from human-written code a cr

MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications

Model ReleasesDGX agent

arXiv:2511.13131v2 Announce Type: replace Abstract: Large Language Models (LLMs) have emerged as powerful tools for automating complex reasoning and decision-making tasks. In telecommunications, they

Nice paper combining the strength of Skills and RAG. Most RAG systems retrieve on every query, whether the model needs help or not. This is …

AgentsDGX agent

Nice paper combining the strength of Skills and RAG. Most RAG systems retrieve on every query, whether the model needs help or not. This is wasteful when the model already knows the answer, and often

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation

Model ReleasesDGX agent

arXiv:2604.15670v1 Announce Type: new Abstract: Reasoning segmentation has recently expanded from ground-level scenes to remote-sensing imagery, yet UAV data poses distinct challenges, including obliq

Puppets or partners? Governing cyborg propaganda in the digital public square

SafetyDGX agent

arXiv:2602.13088v2 Announce Type: replace-cross Abstract: The distinction between genuine grassroots activism and automated influence operations is collapsing. While contemporary policy debates priori

RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models

Model ReleasesDGX agent

arXiv:2601.03699v2 Announce Type: replace Abstract: As large language models (LLMs) become integral to safety-critical applications, ensuring their robustness against adversarial prompts is paramount.

RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity

Model ReleasesDGX agent

arXiv:2509.25897v2 Announce Type: replace-cross Abstract: People often encounter role conflicts -- social dilemmas where the expectations of multiple roles clash and cannot be simultaneously fulfilled

Seed1.8 Model Card: Towards Generalized Real-World Agency

Model ReleasesDGX agent

arXiv:2603.20633v3 Announce Type: replace Abstract: We present Seed1.8, a foundation model aimed at generalized real-world agency: going beyond single-turn prediction to multi-turn interaction, tool u

Sketch and Text Synergy: Fusing Structural Contours and Descriptive Attributes for Fine-Grained Image Retrieval

Model ReleasesDGX agent

arXiv:2604.15735v1 Announce Type: cross Abstract: Fine-grained image retrieval via hand-drawn sketches or textual descriptions remains a critical challenge due to inherent modality gaps. While hand-dr

TabularMath: Understanding Math Reasoning over Tables with Large Language Models

Model ReleasesDGX agent

arXiv:2505.19563v4 Announce Type: replace Abstract: Mathematical reasoning has long been a key benchmark for evaluating large language models. Although substantial progress has been made on math word

Technically Love: The Evolution of Human-AI Romance Discourse on Reddit

ApplicationsDGX agent

arXiv:2604.15333v1 Announce Type: cross Abstract: Human-AI romantic relationships are increasingly common, yet little is understood about how public discourse around them emerges and shifts over time.

The AI industry insists they can manage the risks of superintelligence, but there are in fact zero widely agreed on or accepted solutions to…

SafetyDGX agent

The AI industry insists they can manage the risks of superintelligence, but there are in fact zero widely agreed on or accepted solutions to the problem of how one could even control something vastly

the Codex x @skybysoftware acquisition may have been one of the best @openai deals made in the last year. I've been waiting for 'real' compu…

Model ReleasesDGX agent

the Codex x @skybysoftware acquisition may have been one of the best @openai deals made in the last year. I've been waiting for 'real' computer use since @romainhuet demoed the ChatGPT App with 4o Vis

The Crutch or the Ceiling? How Different Generations of LLMs Shape EFL Student Writings

ApplicationsDGX agent

arXiv:2604.15460v1 Announce Type: cross Abstract: The rapid evolution of Large Language Models (LLMs) has made them powerful tools for enhancing student writing. This study explores the extent and lim

The Jensen + @dwarkesh_sp podcast was fantastic. Jensen is someone who understood how ecosystems work and someone who understands real-world…

Model ReleasesDGX agent

The Jensen + @dwarkesh_sp podcast was fantastic. Jensen is someone who understood how ecosystems work and someone who understands real-world trade, policy and controls work. And in some deeper sense h

The Relic Condition: When Published Scholarship Becomes Material for Its Own Replacement

Model ReleasesDGX agent

arXiv:2604.16116v1 Announce Type: cross Abstract: We extracted the scholarly reasoning systems of two internationally prominent humanities and social science scholars from their published corpora alon

The threat of analytic flexibility in using large language models to simulate human data

HardwareDGX agent

arXiv:2509.13397v3 Announce Type: replace-cross Abstract: Social scientists are now using large language models to create 'silicon samples': synthetic datasets intended to stand in for human responden

Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures

SafetyDGX agent

arXiv:2604.16042v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have achieved strong performance across many NLP tasks, their opaque internal mechanisms hinder trustworthiness and

Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning

SafetyDGX agent

arXiv:2604.15705v1 Announce Type: new Abstract: Reinforcement Fine-Tuning (RFT) has established itself as a critical paradigm for the alignment of Multi-modal Large Language Models (MLLMs) with comple

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik

ToolsDGX agent

Noetik uses transformer neural networks and machine learning to improve cancer drug trial success rates, addressing the historically high failure rate (95%) in clinical development. The approach lever

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions

Model ReleasesDGX agent

arXiv:2604.15823v1 Announce Type: new Abstract: Embodied robotic agents often perceive movies through an egocentric screen-view interface rather than native cinematic footage, introducing domain shift

Weak-Link Optimization for Multi-Agent Reasoning and Collaboration

AgentsDGX agent

arXiv:2604.15972v1 Announce Type: new Abstract: LLM-driven multi-agent frameworks address complex reasoning tasks through multi-role collaboration. However, existing approaches often suffer from reaso

← Previous
1…415416417418419…424
Next →