Freelance data scientist · All case studies
Movie reviews at Map/Reduce scale.
A large rated-review corpus, scored by threshold, with stop words held out through Hadoop’s distributed cache. Python Streaming jobs — not a laptop grep on a sample.
Client: Review corpus study. Built by Dilshad Raza.
Rated reviews at volume do not fit a single-machine count. Score thresholds had to be isolated, and common stop words had to disappear on every mapper — not after a slow reduce of noise.
I implemented Map/Reduce in Python on Hadoop Streaming and pushed the stop-word file through distributed cache so every mapper filtered the same noise. Jobs were run against the rated-review corpus to surface language that actually moved with score.
Hire Dilshad Raza for similar freelance data science, machine learning, and AI automation in the UK, United States, and Australia.