Large Language Models (LLMs) have shown promising capabilities in code generation, yet their application to blockchain smart contracts for high-performance platforms like Solana remains largely unexplored. This paper presents a controlled experimental study evaluating four state-of-the-art LLMs, Claude Sonnet 4.5, GPT-5, Copilot, and DeepSeek V3.2, in generating Solana smart contracts using the Anchor framework. Our research addresses three main objectives: first, assessing the ability of LLMs to generate correct and functional Anchor code; second, exploring prompt engineering techniques to identify successful strategies and common failures; third, classifying hallucinations into syntactic, logical, and interpretive errors using quantitative code metrics. We establish a benchmark of 15 use cases with automated test generation and systematic prompt selection. Results show significant differences across models and prompt strategies, with structured approaches reducing critical errors. Our findings provide practical insights into LLM capabilities and limitations for smart contract generation, enabling better understanding of error patterns and establishing procedures for their detection. This contributes to addressing key challenges in AI-assisted blockchain development within the Solana ecosystem.

Generative Models for Writing Solana Smart Contracts: Comparative Analysis of Performance and Code Metrics

Cuncu, Emanuele;Pinna, Andrea;Ibba, Giacomo Francesco;Baralla, Gavina;Fenu, Gianni;Tonelli, Roberto
2026-01-01

Abstract

Large Language Models (LLMs) have shown promising capabilities in code generation, yet their application to blockchain smart contracts for high-performance platforms like Solana remains largely unexplored. This paper presents a controlled experimental study evaluating four state-of-the-art LLMs, Claude Sonnet 4.5, GPT-5, Copilot, and DeepSeek V3.2, in generating Solana smart contracts using the Anchor framework. Our research addresses three main objectives: first, assessing the ability of LLMs to generate correct and functional Anchor code; second, exploring prompt engineering techniques to identify successful strategies and common failures; third, classifying hallucinations into syntactic, logical, and interpretive errors using quantitative code metrics. We establish a benchmark of 15 use cases with automated test generation and systematic prompt selection. Results show significant differences across models and prompt strategies, with structured approaches reducing critical errors. Our findings provide practical insights into LLM capabilities and limitations for smart contract generation, enabling better understanding of error patterns and establishing procedures for their detection. This contributes to addressing key challenges in AI-assisted blockchain development within the Solana ecosystem.
2026
Code Metrics
Hallucinations
LLM
Prompt Engineering
Smart Contracts
Solana
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I metadati presenti in IRIS UNICA sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono protetti da diritto d'autore, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11584/491346
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 0
social impact