Make a messy assignment export usable
Clean data with an auditable set of rules
← All assignments
Assignment preview
This is the complete public sample: the brief, source notes, deliverables, checks, and review criteria below. Download the same materials to work offline. Confirmed paid assignment terms and submission tools are available only in the invited portal.
The brief
Clean the supplied fictional assignment export. Trim spaces, normalize department names, remove exact duplicate rows after normalization, and quarantine invalid records. Do not guess missing or invalid values.
Your source packet
Fictional source notes for this exercise. No outside client data.
CSV input: id,department,minutes,score A01,Research,30,80 A02, data ,45,90 A02,Data,45,90 A03,Development,-5,70 A04,Quantitative,25, A05,research,40,105 A06,Data,20,60 Valid departments: Research, Quantitative, Development, Data. Minutes must be positive. Scores must be numeric from 0 to 100 inclusive. A missing score goes to quarantine.
What to deliver
- clean.csv containing valid, unique rows
- quarantine.csv with a reason for each rejected row
- A 200–300 word cleaning log, including the duplicate count
Check your understanding
- Which two rows become duplicates after normalization?
- Why should a missing score stay missing instead of becoming zero?