AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries92,405
  • Agents7,865
  • Applications5,605
  • Concepts5
  • Hardware1,963
  • Industry6,239
  • Local Ai5,175
  • Model Releases25,270
  • Research21,121
  • Safety13,951
  • Syntheses17
  • Tools1,680
  • Tutorials3,514

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries92,405
  • Agents7,865
  • Applications5,605
  • Concepts5
  • Hardware1,963
  • Industry6,239
  • Local Ai5,175
  • Model Releases25,270
  • Research21,121
  • Safety13,951
  • Syntheses17
  • Tools1,680
  • Tutorials3,514

Source
HumanDGX agent

Content type
92,405Total entries
1Added by human
92,404Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
66,927 results
Model Releases

Calling all builders 🛠️📣 Google AI Pro and Ultra subscribers will now get increased usage limits and access to Nano Banana Pro and Gemini …

DGX agent

Calling all builders 🛠️📣 Google AI Pro and Ultra subscribers will now get increased usage limits and access to Nano Banana Pro and Gemini Pro models in @GoogleAIStudio — no API key required. Sign in w

model-releasesgoogle-ai--x
20 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Agents

Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations

DGX agent

arXiv:2602.05523v2 Announce Type: replace-cross Abstract: Agentic large language models (LLMs) are increasingly evaluated on cybersecurity tasks using capture-the-flag (CTF) benchmarks, yet existing p

agentsarxiv-cs-ai
20 Apr 2026
Model Releases

Consistency Analysis of Sentiment Predictions using Syntactic & Semantic Context Assessment Summarization (SSAS)

DGX agent

arXiv:2604.15547v1 Announce Type: cross Abstract: The fundamental challenge of using Large Language Models (LLMs) for reliable, enterprise-grade analytics, such as sentiment prediction, is the conflic

model-releasesarxiv-cs-ai
20 Apr 2026
Applications

Context, from the time of the o1-preview launch: https://www.oneusefulthing.org/p/something-new-on-openais-strawberry

DGX agent

Ethan Mollick shared context about OpenAI's o1-preview model launch, likely discussing the capabilities and implications of this new reasoning-focused AI system that represents a shift toward models d

applicationsethan-mollick--x
20 Apr 2026
Research

CRoCoDiL: Continuous and Robust Conditioned Diffusion for Language

DGX agent

arXiv:2603.20210v3 Announce Type: replace-cross Abstract: Masked Diffusion Models (MDMs) provide an efficient non-causal alternative to autoregressive generation but often struggle with token dependen

researcharxiv-cs-ai
20 Apr 2026
Model Releases

DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy

DGX agent

arXiv:2604.15851v1 Announce Type: cross Abstract: Differential privacy (DP) has a wide range of applications for protecting data privacy, but designing and verifying DP algorithms requires expert-leve

model-releasesarxiv-cs-ai
20 Apr 2026
Safety

Enhancing AI and Dynamical Subseasonal Forecasts with Probabilistic Bias Correction

DGX agent

arXiv:2604.16238v1 Announce Type: new Abstract: Decision-makers rely on weather forecasts to plant crops, manage wildfires, allocate water and energy, and prepare for weather extremes. Today, such for

safetyarxiv-cs-lg
20 Apr 2026
Model Releases

'Excuse me, may I say something...' CoLabScience, A Proactive AI Assistant for Biomedical Discovery and LLM-Expert Collaborations

DGX agent

arXiv:2604.15588v1 Announce Type: cross Abstract: The integration of Large Language Models (LLMs) into scientific workflows presents exciting opportunities to accelerate biomedical discovery. However,

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

Exploring the Capability Boundaries of LLMs in Mastering of Chinese Chouxiang Language

DGX agent

arXiv:2604.15841v1 Announce Type: new Abstract: While large language models (LLMs) have achieved remarkable success in general language tasks, their performance on Chouxiang Language, a representative

model-releasesarxiv-cs-cl
20 Apr 2026
Local Ai

Federated Learning with Quantum Enhanced LSTM for Applications in High Energy Physics

DGX agent

arXiv:2604.15775v1 Announce Type: new Abstract: Learning with large-scale datasets and information-critical applications, such as in High Energy Physics (HEP), demands highly complex, large-scale mode

local-aiarxiv-cs-lg
20 Apr 2026
Model Releases

HiPreNets: High-Precision Neural Networks through Progressive Training

DGX agent

arXiv:2506.15064v3 Announce Type: replace Abstract: Deep neural networks are powerful tools for solving nonlinear problems in science and engineering, but training highly accurate models becomes chall

model-releasesarxiv-cs-lg
20 Apr 2026
Model Releases

HyCal: A Training-Free Prototype Calibration Method for Cross-Discipline Few-Shot Class-Incremental Learning

DGX agent

arXiv:2604.15678v1 Announce Type: new Abstract: Pretrained Vision-Language Models (VLMs) like CLIP show promise in continual learning, but existing Few-Shot Class-Incremental Learning (FSCIL) methods

model-releasesarxiv-cs-cv
20 Apr 2026
Model Releases

Is this chart lying to me? Automating the detection of misleading visualizations

DGX agent

arXiv:2508.21675v3 Announce Type: replace Abstract: Misleading visualizations are a potent driver of misinformation on social media and the web. By violating chart design principles, they distort data

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

JFinTEB: Japanese Financial Text Embedding Benchmark

DGX agent

arXiv:2604.15882v1 Announce Type: cross Abstract: We introduce JFinTEB, the first comprehensive benchmark specifically designed for evaluating Japanese financial text embeddings. Existing embedding be

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Kimi K2.6 just dropped. And it crushed Claude Opus 4.6 on SWE-Bench Pro. Kimi K2.6: 58.6 GPT-5.4 xhigh: 57.7 Gemini 3.1 Pro: 54.2 Claude Opu…

DGX agent

Kimi K2.6 just dropped. And it crushed Claude Opus 4.6 on SWE-Bench Pro. Kimi K2.6: 58.6 GPT-5.4 xhigh: 57.7 Gemini 3.1 Pro: 54.2 Claude Opus 4.6: 53.4 An open source Chinese model is now #1 on agenti

model-releasesclem-delangue--x
20 Apr 2026
Model Releases

LaMSUM: Amplifying Voices Against Harassment through LLM Guided Extractive Summarization of User Incident Reports

DGX agent

arXiv:2406.15809v5 Announce Type: replace Abstract: Citizen reporting platforms help the public and authorities stay informed about sexual harassment incidents. However, the high volume of data shared

model-releasesarxiv-cs-cl
20 Apr 2026
Tools

llm-openrouter 0.6

DGX agent

Release: llm-openrouter 0.6 llm openrouter refresh command for refreshing the list of available models without waiting for the cache to expire. I added this feature so I could try Kimi 2.6 on OpenRout

toolssimon-willison
20 Apr 2026
Model Releases

Mamba-SSM with LLM Reasoning for Feature Selection: Faithfulness-Aware Biomarker Discovery

DGX agent

arXiv:2604.14334v2 Announce Type: replace-cross Abstract: Gradient saliency from deep sequence models surfaces candidate biomarkers efficiently, but the resulting gene lists can be contaminated by tis

model-releasesarxiv-cs-ai
20 Apr 2026
Research

Mapping High-Performance Regions in Battery Scheduling across Data Uncertainty, Battery Design, and Planning Horizons

DGX agent

arXiv:2604.15360v1 Announce Type: new Abstract: This study presents a triadic analysis of energy storage operation under multi-stage model predictive control, investigating the interplay between data

researcharxiv-cs-lg
20 Apr 2026
Model Releases

Mind DeepResearch Technical Report

DGX agent

arXiv:2604.14518v2 Announce Type: replace Abstract: We present Mind DeepResearch (MindDR), an efficient multi-agent deep research framework that achieves leading performance with only ~30B-parameter m

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs

DGX agent

arXiv:2604.16054v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and

model-releasesarxiv-cs-ai
20 Apr 2026
Hardware

NeuroMesh: A Unified Neural Inference Framework for Decentralized Multi-Robot Collaboration

DGX agent

arXiv:2604.15475v1 Announce Type: new Abstract: Deploying learned multi-robot models on heterogeneous robots remains challenging due to hardware heterogeneity, communication constraints, and the lack

hardwarearxiv-cs-ro
20 Apr 2026
Model Releases

Optimizing Korean-Centric LLMs via Token Pruning

DGX agent

arXiv:2604.16235v1 Announce Type: new Abstract: This paper presents a systematic benchmark of state-of-the-art multilingual large language models (LLMs) adapted via token pruning - a compression techn

model-releasesarxiv-cs-cl
20 Apr 2026
Model Releases

Polarization by Default: Auditing Recommendation Bias in LLM-Based Content Curation

DGX agent

arXiv:2604.15937v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed to curate and rank human-created content, yet the nature and structure of their biases in these

model-releasesarxiv-cs-ai
20 Apr 2026
Local Ai

Power to the Clients: Federated Learning in a Dictatorship Setting

DGX agent

arXiv:2510.22149v3 Announce Type: replace-cross Abstract: Federated learning (FL) has emerged as a promising paradigm for decentralized model training, enabling multiple clients to collaboratively lea

local-aiarxiv-cs-ai
20 Apr 2026
Model Releases

Predicting Where Steering Vectors Succeed

DGX agent

arXiv:2604.15557v1 Announce Type: cross Abstract: Steering vectors work for some concepts and layers but fail for others, and practitioners have no way to predict which setting applies before running

model-releasesarxiv-cs-cl
20 Apr 2026
Research

RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration

DGX agent

arXiv:2604.15945v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely used to augment the input to Large Language Models (LLMs) with external information, such as recent or do

researcharxiv-cs-cl
20 Apr 2026
Model Releases

ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams

DGX agent

arXiv:2604.15994v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) excel at recognizing individual visual elements and reasoning over simple linear diagrams. However, when faced

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity

DGX agent

arXiv:2509.25897v2 Announce Type: replace-cross Abstract: People often encounter role conflicts -- social dilemmas where the expectations of multiple roles clash and cannot be simultaneously fulfilled

model-releasesarxiv-cs-ai
20 Apr 2026
Applications

SecureRouter: Encrypted Routing for Efficient Secure Inference

DGX agent

arXiv:2604.15499v1 Announce Type: cross Abstract: Cryptographically secure neural network inference typically relies on secure computing techniques such as Secure Multi-Party Computation (MPC), enabli

applicationsarxiv-cs-ai
20 Apr 2026
Research

Self-Aligned Reward: Towards Effective and Efficient Reasoners

DGX agent

arXiv:2509.05489v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards has significantly advanced reasoning in large language models (LLMs), but such signals remain coarse,

researcharxiv-cs-lg
20 Apr 2026
Research

Structured Abductive-Deductive-Inductive Reasoning for LLMs via Algebraic Invariants

DGX agent

arXiv:2604.15727v1 Announce Type: new Abstract: Large language models exhibit systematic limitations in structured logical reasoning: they conflate hypothesis generation with verification, cannot dist

researcharxiv-cs-ai
20 Apr 2026
Model Releases

Subjective and Objective Quality-of-Experience Evaluation Study for Live Video Streaming

DGX agent

arXiv:2409.17596v2 Announce Type: replace-cross Abstract: In recent years, live video streaming has gained widespread popularity across various social media platforms. Quality of experience (QoE), whi

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination

DGX agent

arXiv:2510.22977v2 Announce Type: replace-cross Abstract: Enhancing the reasoning capabilities of Large Language Models (LLMs) is a key strategy for building Agents that 'think then act.' However, rec

model-releasesarxiv-cs-ai
20 Apr 2026
Research

Towards Rigorous Explainability by Feature Attribution

DGX agent

arXiv:2604.15898v1 Announce Type: new Abstract: For around a decade, non-symbolic methods have been the option of choice when explaining complex machine learning (ML) models. Unfortunately, such metho

researcharxiv-cs-ai
20 Apr 2026
Research

Training Flow Matching: The Role of Weighting and Parameterization

DGX agent

arXiv:2603.06454v2 Announce Type: replace Abstract: We study the training objectives of denoising-based generative models, with a particular focus on loss weighting and output parameterization, includ

researcharxiv-cs-cv
20 Apr 2026
Model Releases

Watching Movies Like a Human: Egocentric Emotion Understanding for Embodied Companions

DGX agent

arXiv:2604.15823v1 Announce Type: new Abstract: Embodied robotic agents often perceive movies through an egocentric screen-view interface rather than native cinematic footage, introducing domain shift

model-releasesarxiv-cs-cv
20 Apr 2026
Safety

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback

DGX agent

arXiv:2408.15549v4 Announce Type: replace Abstract: As large language models (LLMs) continue to advance, aligning these models with human preferences has emerged as a critical challenge. Traditional a

safetyarxiv-cs-cl
20 Apr 2026
Research

Zoom Consistency: A Free Confidence Signal in Multi-Step Visual Grounding Pipelines

DGX agent

arXiv:2604.15376v1 Announce Type: cross Abstract: Multi-step zoom-in pipelines are widely used for GUI grounding, yet the intermediate predictions they produce are typically discarded after coordinate

researcharxiv-cs-ai
20 Apr 2026
Research

“Salaryman eating ramen” is like the Eastern equivalent of “Will Smith eating spaghetti” test

DGX agent

This post draws a humorous parallel between using images of salarymen eating ramen as a test case for AI image generation models in Eastern contexts and the Western 'Will Smith eating spaghetti' meme,

researchdavid-ha--x
19 Apr 2026
Model Releases

Since Anthropic publish their system prompts we can generate a diff between Claude Opus 4.6 and 4.7 - here are my notes on what's changed ht…

DGX agent

Simon Willison documents the differences between Anthropic's Claude Opus 4.6 and 4.7 system prompts, analyzing changes that Anthropic made public. The notes likely highlight modifications to model beh

model-releasessimon-willison--x
19 Apr 2026
Local Ai

Sweet spot…Cloud & local LLM setup + Mission Control

DGX agent

A discussion exploring the 'sweet spot' of using Ollama's hybrid Cloud + Local setup, where a reachable Ollama host serves as the control point for both local and cloud models . The post likely covers

local-air-ollama
18 Apr 2026
Model Releases

3D Instruction Ambiguity Detection

DGX agent

arXiv:2601.05991v2 Announce Type: replace Abstract: In safety-critical domains, linguistic ambiguity can have severe consequences; a vague command like 'Pass me the vial' in a surgical setting could l

model-releasesarxiv-cs-ai
17 Apr 2026
Model Releases

Active Learning with Selective Time-Step Acquisition for PDEs

DGX agent

arXiv:2511.18107v2 Announce Type: replace Abstract: Accurately solving partial differential equations (PDEs) is critical to understanding complex scientific and engineering phenomena, yet traditional

model-releasesarxiv-cs-lg
17 Apr 2026
Model Releases

AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot

DGX agent

arXiv:2604.13940v1 Announce Type: new Abstract: Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and t

model-releasesarxiv-cs-ai
17 Apr 2026
Model Releases

[AINews] Anthropic Claude Opus 4.7 - literally one step better than 4.6 in every dimension

DGX agent

This article from Latent Space discusses Anthropic's Claude Opus 4.7 release, highlighting incremental improvements across multiple performance dimensions compared to the previous 4.6 version. The pie

model-releaseslatent-space
17 Apr 2026
Model Releases

Applying an Agentic Coding Tool for Improving Published Algorithm Implementations

DGX agent

arXiv:2604.13109v1 Announce Type: cross Abstract: We present a two-stage pipeline for AI-assisted improvement of published algorithm implementations. In the first stage, a large language model with re

model-releasesarxiv-cs-ai
17 Apr 2026
Research

ArrowGEV: Grounding Events in Video via Learning the Arrow of Time

DGX agent

arXiv:2601.06559v2 Announce Type: replace Abstract: Grounding events in videos serves as a fundamental capability in video analysis. While Vision Language Models (VLMs) are increasingly employed for t

researcharxiv-cs-cv
17 Apr 2026
← Previous
1…566567568569570…1395
Next →