Make a messy assignment export usable

Clean data with an auditable set of rules

← All assignments

Assignment preview

This is the complete public sample: the brief, source notes, deliverables, checks, and review criteria below. Download the same materials to work offline. Confirmed paid assignment terms and submission tools are available only in the invited portal.

The brief

Clean the supplied fictional assignment export. Trim spaces, normalize department names, remove exact duplicate rows after normalization, and quarantine invalid records. Do not guess missing or invalid values.

Your source packet

Fictional source notes for this exercise. No outside client data.

CSV input:
id,department,minutes,score
A01,Research,30,80
A02, data ,45,90
A02,Data,45,90
A03,Development,-5,70
A04,Quantitative,25,
A05,research,40,105
A06,Data,20,60

Valid departments: Research, Quantitative, Development, Data. Minutes must be positive. Scores must be numeric from 0 to 100 inclusive. A missing score goes to quarantine.

What to deliver

  1. clean.csv containing valid, unique rows
  2. quarantine.csv with a reason for each rejected row
  3. A 200–300 word cleaning log, including the duplicate count

Check your understanding

  1. Which two rows become duplicates after normalization?
  2. Why should a missing score stay missing instead of becoming zero?

Request a beta invitation

Request early access