Large Language Models (LLMs) have shown promising capabilities in code generation, yet their application to blockchain smart contracts for high-performance platforms like Solana remains largely unexplored. This paper presents a controlled experimental study evaluating four state-of-the-art LLMs, Claude Sonnet 4.5, GPT-5, Copilot, and DeepSeek V3.2, in generating Solana smart contracts using the Anchor framework. Our research addresses three main objectives: first, assessing the ability of LLMs to generate correct and functional Anchor code; second, exploring prompt engineering techniques to identify successful strategies and common failures; third, classifying hallucinations into syntactic, logical, and interpretive errors using quantitative code metrics. We establish a benchmark of 15 use cases with automated test generation and systematic prompt selection. Results show significant differences across models and prompt strategies, with structured approaches reducing critical errors. Our findings provide practical insights into LLM capabilities and limitations for smart contract generation, enabling better understanding of error patterns and establishing procedures for their detection. This contributes to addressing key challenges in AI-assisted blockchain development within the Solana ecosystem.
Generative Models for Writing Solana Smart Contracts: Comparative Analysis of Performance and Code Metrics
Cuncu, Emanuele;Pinna, Andrea;Ibba, Giacomo Francesco;Baralla, Gavina;Fenu, Gianni;Tonelli, Roberto
2026-01-01
Abstract
Large Language Models (LLMs) have shown promising capabilities in code generation, yet their application to blockchain smart contracts for high-performance platforms like Solana remains largely unexplored. This paper presents a controlled experimental study evaluating four state-of-the-art LLMs, Claude Sonnet 4.5, GPT-5, Copilot, and DeepSeek V3.2, in generating Solana smart contracts using the Anchor framework. Our research addresses three main objectives: first, assessing the ability of LLMs to generate correct and functional Anchor code; second, exploring prompt engineering techniques to identify successful strategies and common failures; third, classifying hallucinations into syntactic, logical, and interpretive errors using quantitative code metrics. We establish a benchmark of 15 use cases with automated test generation and systematic prompt selection. Results show significant differences across models and prompt strategies, with structured approaches reducing critical errors. Our findings provide practical insights into LLM capabilities and limitations for smart contract generation, enabling better understanding of error patterns and establishing procedures for their detection. This contributes to addressing key challenges in AI-assisted blockchain development within the Solana ecosystem.I metadati presenti in IRIS UNICA sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono protetti da diritto d'autore, salvo diversa indicazione.



