AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,832
  • Agents7,214
  • Applications5,155
  • Concepts5
  • Hardware1,742
  • Industry6,086
  • Local Ai4,673
  • Model Releases22,315
  • Research19,015
  • Safety12,707
  • Syntheses17
  • Tools1,664
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,832Total entries
1Added by human
83,831Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,323 results
Model Releases

Frequency-Enhanced Dual-Subspace Networks for Few-Shot Fine-Grained Image Classification

DGX agent

arXiv:2604.14958v1 Announce Type: new Abstract: Few-shot fine-grained image classification aims to recognize subcategories with high visual similarity using only a limited number of annotated samples.

model-releasesarxiv-cs-cv
17 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

From Memorization to Creativity: LLM as a Designer of Novel Neural Architectures

DGX agent

arXiv:2601.02997v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel in program synthesis, yet their capacity for neural architecture design -- balancing syntactic reliability,

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis

DGX agent

arXiv:2604.13888v1 Announce Type: new Abstract: The integration of Large Language Models (LLMs) into Geographic Information Systems (GIS) marks a paradigm shift toward autonomous spatial analysis. How

model-releasesarxiv-cs-ai
17 Apr 2026
Model Releases

GraphScout: Empowering Large Language Models with Intrinsic Exploration Ability for Agentic Graph Reasoning

DGX agent

arXiv:2603.01410v2 Announce Type: replace Abstract: Knowledge graphs provide structured and reliable information for many real-world applications, motivating increasing interest in combining large lan

model-releasesarxiv-cs-ai
17 Apr 2026
Model Releases

H2VLR: Heterogeneous Hypergraph Vision-Language Reasoning for Few-Shot Anomaly Detection

DGX agent

arXiv:2604.14507v1 Announce Type: new Abstract: As a classic vision task, anomaly detection has been widely applied in industrial inspection and medical imaging. In this task, data scarcity is often a

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

Hierarchical vs. Flat Iteration in Shared-Weight Transformers

DGX agent

arXiv:2604.14442v1 Announce Type: new Abstract: We present an empirical study of whether hierarchically structured, shared-weight recurrence can match the representational quality of independent-layer

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

How Embeddings Shape Graph Neural Networks: Classical vs Quantum-Oriented Node Representations

DGX agent

arXiv:2604.15273v1 Announce Type: new Abstract: Node embeddings act as the information interface for graph neural networks, yet their empirical impact is often reported under mismatched backbones, spl

model-releasesarxiv-cs-lg
17 Apr 2026
Model Releases

HRDexDB: A Large-Scale Dataset of Dexterous Human and Robotic Hand Grasps

DGX agent

arXiv:2604.14944v1 Announce Type: cross Abstract: We present HRDexDB, a large-scale, multi-modal dataset of high-fidelity dexterous grasping sequences featuring both human and diverse robotic hands. U

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

I have found 4.7 great for design, reverted back to 4.6 extended for everything else Anyone else like this?

DGX agent

I have found 4.7 great for design, reverted back to 4.6 extended for everything else Anyone else like this? Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talk

model-releasesemad-mostaque--x
17 Apr 2026
Model Releases

I was told by Anthropic that they are looking at ways of fixing this, which is good (you can also see a reply from a Claude PM in the thread…

DGX agent

I was told by Anthropic that they are looking at ways of fixing this, which is good (you can also see a reply from a Claude PM in the thread). I think the adaptive thinking requirement in Claude Opus

model-releasesethan-mollick--x
17 Apr 2026
Model Releases

IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation

DGX agent

arXiv:2511.01014v3 Announce Type: replace Abstract: Instruction-following is a fundamental ability of Large Language Models (LLMs), requiring their generated outputs to follow multiple constraints imp

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation

DGX agent

arXiv:2603.04738v2 Announce Type: replace Abstract: Instruction-following is a foundational capability of large language models (LLMs), with its improvement hinging on scalable and accurate feedback f

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models

DGX agent

arXiv:2604.08064v2 Announce Type: replace Abstract: Existing memory benchmarks for LLM agents evaluate explicit recall of facts, yet overlook implicit memory where experience becomes automated behavio

model-releasesarxiv-cs-ai
17 Apr 2026
Model Releases

Improving Clean Accuracy via a Tangent-Space Perspective on Adversarial Training

DGX agent

arXiv:2408.14728v2 Announce Type: replace Abstract: Adversarial training has proven effective in improving the robustness of deep neural networks against adversarial attacks. However, this enhanced ro

model-releasesarxiv-cs-lg
17 Apr 2026
Model Releases

Improving Language Models with Intentional Analysis

DGX agent

arXiv:2502.04689v4 Announce Type: replace Abstract: Intent, a critical cognitive notion and mental state, is ubiquitous in human communication and problem-solving. Accurately understanding the underly

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

In Context Learning and Reasoning for Symbolic Regression with Large Language Models

DGX agent

arXiv:2410.17448v3 Announce Type: replace Abstract: Large Language Models (LLMs) are transformer-based machine learning models that have shown remarkable performance in tasks for which they were not e

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model

DGX agent

arXiv:2604.14180v1 Announce Type: new Abstract: We train a 318M-parameter Transformer language model from scratch on a curated corpus of 1.56 billion tokens of pure Classical Chinese, with zero Englis

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude. Powered by Claude Opus 4.7, our m…

DGX agent

Introducing Claude Design by Anthropic Labs: make prototypes, slides, and one-pagers by talking to Claude. Powered by Claude Opus 4.7, our most capable vision model. Available in research preview on t

model-releasesthariq--x
17 Apr 2026
Model Releases

Join us at PyCon US 2026 in Long Beach - we have new AI and security tracks this year

DGX agent

This year's PyCon US is coming up next month from May 13th to May 19th, with the core conference talks from Friday 15th to Sunday 17th and tutorial and sprint days either side. It's in Long Beach, Cal

model-releasessimon-willison
17 Apr 2026
Model Releases

Knowing When Not to Answer: Evaluating Abstention in Multimodal Reasoning Systems

DGX agent

arXiv:2604.14799v1 Announce Type: new Abstract: Effective abstention (EA), recognizing evidence insufficiency and refraining from answering, is critical for reliable multimodal systems. Yet existing e

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Label-efficient underwater species classification with logistic regression on frozen foundation model embeddings

DGX agent

arXiv:2604.00313v2 Announce Type: replace Abstract: Automated species classification from underwater imagery is bottlenecked by the cost of expert annotation, and supervised models trained on one data

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

Last week, Anthropic announced Project Glasswing alongside Claude Mythos Preview, a model they described as so powerful at finding vulnerabi…

DGX agent

Last week, Anthropic announced Project Glasswing alongside Claude Mythos Preview, a model they described as so powerful at finding vulnerabilities they couldn't release it. The announcement featured A

model-releasesemad-mostaque--x
17 Apr 2026
Model Releases

Learning Where to Embed: Noise-Aware Positional Embedding for Query Retrieval in Small-Object Detection

DGX agent

arXiv:2604.15065v1 Announce Type: new Abstract: Transformer-based detectors have advanced small-object detection, but they often remain inefficient and vulnerable to background-induced query noise, wh

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

Let's talk content faithfulness. Four days ago, we launched ParseBench, the first document OCR benchmark for AI agents. Its most fundamental…

DGX agent

Let's talk content faithfulness. Four days ago, we launched ParseBench, the first document OCR benchmark for AI agents. Its most fundamental metric asks: did the parser capture all the text, in order,

model-releasesjerry-liu--x
17 Apr 2026
Model Releases

LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligence

DGX agent

arXiv:2512.04578v3 Announce Type: replace Abstract: Legal general intelligence (GI) refers to artificial intelligence (AI) that encompasses legal understanding, reasoning, and decision-making, simulat

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Listen, Correct, and Feed Back: Spoken Pedagogical Feedback Generation

DGX agent

arXiv:2604.14177v1 Announce Type: new Abstract: Grammatical error correction (GEC) and explanation (GEE) have made rapid progress, but real teaching scenarios also require learner-friendly pedagogical

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Listen to the OpenAI Podcast on— Spotify https://open.spotify.com/show/0zojMEDizKMh3aTxnGLENP Apple https://podcasts.apple.com/us/podcast/op…

DGX agent

Listen to the OpenAI Podcast on— Spotify https://open.spotify.com/show/0zojMEDizKMh3aTxnGLENP Apple https://podcasts.apple.com/us/podcast/openai-podcast/id1820330260 YouTube https://youtu.be/UZyH0nx5z

model-releasesopenai--x
17 Apr 2026
Model Releases

LLM Predictive Scoring and Validation: Inferring Experience Ratings from Unstructured Text

DGX agent

arXiv:2604.14321v1 Announce Type: new Abstract: We tasked GPT-4.1 to read what baseball fans wrote about their game-day experience and predict the overall experience rating each fan gave on a 0-10 sur

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems

DGX agent

arXiv:2601.14053v2 Announce Type: replace-cross Abstract: The field of artificial intelligence has undergone a revolution from foundational Transformer architectures to reasoning-capable systems appro

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

DGX agent

arXiv:2604.15149v1 Announce Type: new Abstract: As reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for scaling reasoning capabilities in LLMs, a new failure mode

model-releasesarxiv-cs-lg
17 Apr 2026
Model Releases

LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning

DGX agent

arXiv:2604.14922v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has emerged as a critical driver for enhancing the reasoning capabilities of Large Language Models (LLMs). While recent ad

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

MADE: A Living Benchmark for Multi-Label Text Classification with Uncertainty Quantification of Medical Device Adverse Events

DGX agent

arXiv:2604.15203v1 Announce Type: new Abstract: Machine learning in high-stakes domains such as healthcare requires not only strong predictive performance but also reliable uncertainty quantification

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Magnitude Is All You Need? Rethinking Phase in Quantum Encoding of Complex SAR Data

DGX agent

arXiv:2604.14229v1 Announce Type: cross Abstract: Synthetic Aperture Radar (SAR) data is inherently complex-valued, while quantum machine learning (QML) models naturally operate in complex Hilbert spa

model-releasesarxiv-cs-lg
17 Apr 2026
Model Releases

MARCA: A Checklist-Based Benchmark for Multilingual Web Search

DGX agent

arXiv:2604.14448v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select rel

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents

DGX agent

arXiv:2509.06477v2 Announce Type: replace Abstract: Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, fostering a promising hybrid paradigm for ML

model-releasesarxiv-cs-ai
17 Apr 2026
Model Releases

Maximal Brain Damage Without Data or Optimization: Disrupting Neural Networks via Sign-Bit Flips

DGX agent

arXiv:2502.07408v2 Announce Type: replace-cross Abstract: Deep Neural Networks (DNNs) can be catastrophically disrupted by flipping only a handful of parameter bits. We introduce Deep Neural Lesion (D

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

Measuring multi-calibration

DGX agent

arXiv:2506.11251v2 Announce Type: replace-cross Abstract: A suitable scalar metric can help measure multi-calibration, defined as follows. When the expected values of observed responses are equal to c

model-releasesarxiv-cs-lg
17 Apr 2026
Model Releases

MEBench: A Novel Benchmark for Understanding Mutual Exclusivity Bias in Vision-Language Models

DGX agent

arXiv:2505.20122v2 Announce Type: replace Abstract: This paper introduces MEBench, a novel benchmark for evaluating mutual exclusivity (ME) bias, a cognitive phenomenon observed in children during wor

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

Mechanistic Decoding of Cognitive Constructs in LLMs

DGX agent

arXiv:2604.14593v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate increasingly sophisticated affective capabilities, the internal mechanisms by which they process complex

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

MemGround: Long-Term Memory Evaluation Kit for Large Language Models in Gamified Scenarios

DGX agent

arXiv:2604.14158v1 Announce Type: new Abstract: Current evaluations of long-term memory in LLMs are fundamentally static. By fixating on simple retrieval and short-context inference, they neglect the

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

MetaDent: Labeling Clinical Images for Vision-Language Models in Dentistry

DGX agent

arXiv:2604.14866v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated significant potential in medical image analysis, yet their application in intraoral photography remains

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

Mitigating LLM biases toward spurious social contexts using direct preference optimization

DGX agent

arXiv:2604.02585v2 Announce Type: replace-cross Abstract: LLMs are increasingly used for high-stakes decision-making, yet their sensitivity to spurious contextual information can introduce harmful bia

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining

DGX agent

arXiv:2604.14198v1 Announce Type: cross Abstract: Domain reweighting can improve sample efficiency and downstream generalization, but data-mixture optimization for multimodal midtraining remains large

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation

DGX agent

arXiv:2604.15309v1 Announce Type: cross Abstract: The rapid progress of Artificial Intelligence Generated Content (AIGC) tools enables images, videos, and visualizations to be created on demand for we

model-releasesarxiv-cs-cl
17 Apr 2026
Model Releases

Modular Continual Learning via Zero-Leakage Reconstruction Routing and Autonomous Task Discovery

DGX agent

arXiv:2604.14375v1 Announce Type: new Abstract: Catastrophic forgetting remains a primary hurdle in sequential task learning for artificial neural networks. We propose a silicon-native modular archite

model-releasesarxiv-cs-lg
17 Apr 2026
Model Releases

ModuSeg: Decoupling Object Discovery and Semantic Retrieval for Training-Free Weakly Supervised Segmentation

DGX agent

arXiv:2604.07021v2 Announce Type: replace Abstract: Weakly supervised semantic segmentation aims to achieve pixel-level predictions using image-level labels. Existing methods typically entangle semant

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

Multigrain-aware Semantic Prototype Scanning and Tri-Token Prompt Learning Embraced High-Order RWKV for Pan-Sharpening

DGX agent

arXiv:2604.14622v1 Announce Type: new Abstract: In this work, we propose a Multigrain-aware Semantic Prototype Scanning paradigm for pan-sharpening, built upon a high-order RWKV architecture and a tri

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

Neuro-Oracle: A Trajectory-Aware Agentic RAG Framework for Interpretable Epilepsy Surgical Prognosis

DGX agent

arXiv:2604.14216v1 Announce Type: cross Abstract: Predicting post-surgical seizure outcomes in pharmacoresistant epilepsy is a clinical challenge. Conventional deep-learning approaches operate on stat

model-releasesarxiv-cs-cl
17 Apr 2026
← Previous
1…423424425426427…466
Next →