Freelance data scientist · All case studies

Propaganda is a span, not a document label.

Two sequence-labelling approaches on the same text: a BERT token-classification head with BIO tags (B-Loaded_Language and the rest) via Hugging Face, and SpanMarker to pull the exact propaganda spans out of the table.

Client: Media-integrity NLP. Built by Dilshad Raza.

Loaded language, name-calling, and the rest live inside a sentence. Document classification will call the page dirty and leave the editor guessing where. The work had to mark the span and the technique — BIO on tokens, and a second architecture aimed at exact boundaries.

I trained a BERT token-classification head for standard BIO tags through Hugging Face’s sequence trainer, then evaluated SpanMarker on the same text tables to extract exact propaganda spans. Dual modelling so a missed boundary on one stack was not invisible on the other.

Hire Dilshad Raza for similar freelance data science, machine learning, and AI automation in the UK, United States, and Australia.