Designed by BPER, a partner company of the Foundation, the fifth challenge of the IFAB 4 NEXT GENERATION TALENTS program gives students the opportunity to test themselves with innovative, cutting-edge, and optimized solutions in the field of synthetic data—data that is artificially generated rather than derived from real-world events.
BPER and Synthetic Data: The Challenge
BPER’s challenge requires students to research and implement methodologies for synthetic data generation (both textual and non-textual), using an AI Agent paradigm (consisting of an LLM or LMM, a context, and a prompt). This means developing solutions without training on data, relying solely on zero-shot or few-shot learning approaches.
To meet BPER’s requirements, participants will work with public datasets and follow a structured workflow consisting of multiple steps:
- Selecting the type of data (pure text, PDF documents, tabular data)
- Developing code using Object-Oriented Programming (OOP) in Python
- Documenting the experimentation process and providing conclusions on the work performed
Impact and Benefits
This exciting and forward-thinking challenge will introduce students to synthetic data, enabling them to propose effective usage methodologies. Synthetic data can have significant business applications, particularly in scenarios where:
- Labeling real data is expensive
- Sensitive data cannot be used due to privacy concerns
- Artificial samples are needed to enhance dataset quality and improve the generalization capabilities of algorithms in real-world applications
Through this challenge, students will not only gain hands-on experience in synthetic data generation, but also contribute to developing valuable solutions that can drive innovation in data science and AI applications.
