STARTER MISSION / MACHINE LEARNING
Build a leakage-aware classification experiment
Build a small tabular classifier and show how you evaluated it. Compare a simple baseline, explain the data split and examine errors instead of reporting one attractive score.
The build brief
Work through the requirements in order. Keep your test notes as you go.
- 01
Choose a defensible dataset
Use a small licensed or synthetic dataset. Record the source, permitted use, target and features. Avoid private or identifying records.
- 02
Define the split first
Separate training and evaluation before fitting transformations. If data has time or repeated entities, explain how the split respects that structure.
- 03
Compare and inspect
Use a simple baseline and metrics suited to the task. Record the split, seed, environment, errors and failure patterns.
- 04
Write the model card
Explain intended use, data, evaluation, personal changes, limitations and steps to reproduce the result.
Minimum evidence checklist
These are self-checks. A checked box does not independently verify the work.
- Data source, permission, target and relevant exclusions are documented.
- The split rationale handles time or repeated entities where relevant.
- Preprocessing is fitted using training data only.
- A simple baseline and suitable metrics use the same evaluation setup.
- Error analysis includes examples and limitations, not just a headline score.
- The environment, reproduction steps and model card are included.
An honest project summary
A worked writing example, not a completed reference build.
I compared a simple baseline with a classifier using the same holdout set. Preprocessing was fitted on training data. I recorded errors and the evaluation settings. The experiment is small and has not been validated on real deployment data.
Before you call it done
A useful handoff makes the gaps visible.
- 01
Attach real artifacts
Use your actual repository, test notes, dated screenshots or a setup explanation. Do not invent a result to fill a missing field.
- 02
Explain your contribution
Credit a tutorial, template, dataset or teammate. State what you changed and why.
- 03
Mark what remains unknown
Describe tests you did not run and evidence you cannot safely share. Respect source licenses and agreements.