Fictional resume example. Names, employment histories and results are illustrative, not an actual employee record.
Lewis Clarke
Data Scientist with experience in problem framing, experiments, models, and measurable decisions. Practical work includes Python, SQL, Experimentation, Causal inference.
Experience
Amazon
Seattle · United States
Data Scientist
Mar 2022 - now
- Used SQL and Python to study seller retention, separating new-seller cohorts from established shops and checking seasonality; reported uncertainty rather than treating a correlation as a causal effect.
- Built Looker dashboards from the research dataset with documented cohort definitions and drill-through samples; reconciled dashboard totals to the source queries before sharing recommendations.
- Separated training and evaluation data by time and customer, inspected leakage in derived features, and compared model performance across cohorts before recommending an experiment.
- Built a Python machine learning baseline for repeat-purchase prediction, split observations by customer and time, and compared calibration and recall with a simple statistical baseline before presenting the model recommendation.
- Prepared an experiment readout with confidence intervals, cohort-level diagnostics and a data visualization of the pre-test trend; checked sample-ratio mismatch and explained when the evidence did not justify rollout.
- Owned a forecasting evaluation using time-based splits in Python, reducing holdout error from 19% to 15% against the same seasonal baseline and recording cohort failures.
- Published a Tableau evaluation dashboard using SQL-validated inputs, resolving differences between data-pipeline totals and the statistics presented to decision makers.
Microsoft
Redmond, Washington · United States
Data Scientist
Jan 2019 - Feb 2022
- Built a demand forecast with promotion and holiday effects for 26 regions. Kept the improvement stable through the follow-up review.
- Investigated a discrepancy between a published metric and its source records, traced the transformation that changed the population, and corrected the calculation with a reproducible query.
- Compared the last successful data refresh with a failed run, separated missing source data from transformation errors, and reran only the affected interval. Kept the original query and corrected result together for review.
Selected project
Data Scientist — independent case study
Project owner
Feb 2024 - Jun 2024
- Reconciled three conflicting definitions of active supply into one governed metric
- Generated synthetic source records with duplicates, late arrivals and corrected values; wrote assertions for row counts and key uniqueness, and recorded the expected effect of each case on the reported metric.
- Compared the analytical output with a manually calculated reference table, traced differences to a transformation step, and retained a data dictionary and rerun instructions alongside the corrected query.
- Owned the synthetic-data validation using a manually calculated reference, resolved duplicate-key inflation and completed a notebook that reproduces the corrected totals.
Education
University of Washington
Seattle, Washington · United States
B.S. Computer Science
Sep 2013 - Jun 2017
Relevant coursework: Algorithms, operating systems, databases, computer networks
Skills
Role expertise
Python · SQL · Experimentation · Causal inference · Machine learning · Data visualization
Certifications
Google Advanced Data Analytics Professional Certificate
Jun 2024
Publications
- Published an independent methods note using reproducible queries, explaining the data grain, excluded records and sensitivity of the result to a changed denominator.


