The use of Ethereum smart contracts has significantly influenced sectors that depend on decentralized control and automated financial transactions. However, ensuring their security and reliable execution remains a complex task. Among the most serious challenges is the Denial of Service (DoS) attack, which can make a contract nonfunctional. The broad range of vulnerabilities that enable these attacks complicates prevention efforts. While dynamic security tools exist, they require substantial computational resources, and machine learning-based approaches face limitations due to a lack of training data. To address this issue, we propose a methodology using Large Language Models (LLMs), specifically Antropic's Claude and OpenAI's GPT-4, to generate synthetic examples of Ethereum smart contracts exposed to DoS attacks. Our results show that, with properly designed prompts, these models can produce high-quality synthetic examples, enabling the development of classification and anomaly detection models.
Large language models for synthetic dataset generation: a case study on ethereum smart contract DoS vulnerabilities
Ibba, Giacomo
;Baralla, Gavina
;Destefanis, Giuseppe
2025-01-01
Abstract
The use of Ethereum smart contracts has significantly influenced sectors that depend on decentralized control and automated financial transactions. However, ensuring their security and reliable execution remains a complex task. Among the most serious challenges is the Denial of Service (DoS) attack, which can make a contract nonfunctional. The broad range of vulnerabilities that enable these attacks complicates prevention efforts. While dynamic security tools exist, they require substantial computational resources, and machine learning-based approaches face limitations due to a lack of training data. To address this issue, we propose a methodology using Large Language Models (LLMs), specifically Antropic's Claude and OpenAI's GPT-4, to generate synthetic examples of Ethereum smart contracts exposed to DoS attacks. Our results show that, with properly designed prompts, these models can produce high-quality synthetic examples, enabling the development of classification and anomaly detection models.| File | Dimensione | Formato | |
|---|---|---|---|
|
9_Large_Language_Models_for_Synthetic_Dataset_Generation.pdf
Solo gestori archivio
Descrizione: VoR
Tipologia:
versione editoriale (VoR)
Dimensione
953.2 kB
Formato
Adobe PDF
|
953.2 kB | Adobe PDF | Visualizza/Apri Richiedi una copia |
|
IWBOSE_2025_Iris.pdf
accesso aperto
Descrizione: AAM
Tipologia:
versione post-print (AAM)
Dimensione
957.33 kB
Formato
Adobe PDF
|
957.33 kB | Adobe PDF | Visualizza/Apri |
I metadati presenti in IRIS UNICA sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono protetti da diritto d'autore, salvo diversa indicazione.



