CyberGAN: Generating High-fidelity Cybersecurity Data With Generative Adversarial Networks
Machine learning for cyber defense offers the promise of detecting adversarial activity against the ground data systems managing critical space assets. A fundamental challenge facing machine learning research in cybersecurity is the lack of high-fidelity, shareable datasets for r
Selection note: Curated because “CyberGAN: Generating High-fidelity Cybersecurity Data With Generative Adversarial Networks” covers Machine learning for cyber defense offers the promise of detecting adversarial activity against the ground data systems managing; it materially informs GShips work on ai cybersecurity data.
Evidence boundary: NTRS provides metadata and an abstract, not reviewed full text; methods, results, and current applicability remain unverified. Inclusion is contextual, not automatic claim evidence.
Stable record
ntrs-20220001546
Topic
cybersecurity
Type
Preprint (Draft being sent to journal)
Publisher
Pasadena, CA: Jet Propulsion Laboratory, National Aeronautics and Space Administration, 2020
Authors
Zhang, Yuening; Viswanathan, Arun A; Le, Joie; Gonik, Julia
Year
2020
Editorial state
metadata curated editorial draft
Reviewer
GShips Project editorial synthesis
Official link checked
2026-07-25
Source-supplied abstract
Machine learning for cyber defense offers the promise of detecting adversarial activity against the ground data systems managing critical space assets. A fundamental challenge facing machine learning research in cybersecurity is the lack of high-fidelity, shareable datasets for robust evaluation and testing of machine learning-based solutions. High-fidelity, real-world datasets are necessary for reliable benchmarking of nominal system behavior and malicious activity. Unfortunately, such realistic datasets of both nominal and adversarial activity are rarely shared publicly by data owners due to security and privacy concerns. Besides, the available adversarial data is sparse, which makes training models on malicious activity much harder. This situation has impeded and continues to impede the research and successful adoption of machine learning methods for cyber defense. Researchers have dealt with this problem by generating data within a low-fidelity lab environment, using classified and thus unshareable datasets, or downloading low-fidelity public datasets made available by others. We propose an innovative solution to the problem by employing machine learning methods to generate high-fidelity data. Specifically, we propose the use of Generative Adversarial Networks (GANs) to generate high-fidelity data for cybersecurity purposes. GANs have found successful image processing and natural language applications, but have not yet been investigated for cyber data generation. Our proposed approach first involves training the `discriminator' network of the GAN with a sample of real-world data consisting of malicious and nominal samples. We then use the `generator' network to generate new high-fidelity data samples consisting of an appropriate mix of malicious and nominal activity. We demonstrate applications of our architecture by generating high-fidelity cybersecurity data containing both malicious and nominal samples. We thoroughly evaluate the fidelity of our generated data using heuristics and evaluate its usefulness for machine learning applications using three different datasets. Overall, our approach results in high-fidelity, shareable datasets.
Abstract text has not been adopted as a GShips conclusion.
What would change this record?
A newer or corrected version, a retraction, a verified duplicate, a material topic mismatch, a changed access state, or claim-level review would trigger a dated editorial update.
NTRS provides metadata and an abstract, not reviewed full text; methods, results, and current applicability remain unverified. Inclusion is contextual, not automatic claim evidence.
Source-supplied titles, abstracts, authors, and dates may require correction against the canonical full text.
What would change this page?
A newer or corrected version, retraction, verified duplicate, material topic mismatch, changed access state, or claim-level assessment would change this record.
People, review, and conflicts
Prepared by
GShips Project
Editorial status
metadata-curated-editorial-draft
Editorial reviewer
GShips Project editorial synthesis
Last editorial review
No editorial-review date recorded
Independent review
pending
Independent reviewer
No independent reviewer assigned
Last independent review
No independent-review date exists
Last content edit
Not recorded separately
Official source or link verified
2026-07-25
Declared conflicts
The maintainer intends to explore a commercial venture based on some GShips work. No entity, outside funding, customer, sponsor, or indexed-organization relationship currently exists.