// computer science · applied ai · machine learning · data

Mohona Yesmin.

I’m a Computer Science student at Hunter College building applied AI and machine-learning systems, data pipelines and open-source tools, and I check that a model’s numbers hold up on data it hasn’t seen.

New York, NY  ·  open to roles  ·  Applied AI / ML / data / SWE

  • nowResearch Assistant, DAIR Lab
  • open sourceKubeflow org member
  • buildingSera — adaptive reading app
scroll

01 about

How I got here.

I started in neuroscience, asking how people think and behave. Climate data gave those questions something to measure, from 190+ economies down to single NYC buildings. Now I build machine learning, which grew out of studying how brains learn, and I like turning messy data into one answer I can defend.

education
Hunter College — B.A. Computer Science, minor in Math & Statistics May 2028
Wesleyan University — B.A. Neuroscience & Behavior, minors in Chemistry & Film May 2023
certified
Hugging Face AI Agents · McKinsey Forward Program
leadership
Young Climate Leaders of Color (Fellow) · TEDxWesleyanU (Marketing Intern) · NORD (Vice President)
languages
English, Bengali, Hindi-Urdu native · Mandarin conversational · Spanish basic
The stack, exploded An exploded isometric diagram of five layers: data, pipelines, models, agents and serving, with the tools used at each layer. 01 · data SQL · Supabase · S3 02 · pipelines ETL · scraping · Power BI 03 · models PyTorch · PyG · sklearn 04 · agents LLMs · RAG · MCP 05 · serve FastAPI · Docker · Vercel

02 experience

Where I’ve been building.

  1. 2026

    Research Assistant

    at DAIR Lab

    Aug 2026 – Present
    New York, NY

    • Validated and scaled a multimodal graph neural network pipeline (PyTorch Geometric GAT/GCN encoders, ESM protein embeddings) for toxin ion-channel classification, reaching 87.5% test accuracy on a 200-protein benchmark.
    • Debugged a label-mapping defect that was silently corrupting class-level statistics, and extended a persistent-homology analysis (GUDHI) from a 3-sample proof of concept to the full benchmark.
    • Implementing a custom Gradient Reversal Layer (PyTorch autograd.Function) for domain-adversarial training, to close a ~30-point generalization gap between random-split (~86%) and held-out-taxon (~55%) accuracy.
    • PyTorch
    • PyTorch Geometric
    • GNNs
    • GUDHI
    • ESM embeddings
  2. 2025

    Energy Data Intern → Energy Data Manager I

    at EN-POWER Group

    Jun 2025 – Present
    New York, NY

    • Developed and deployed statistical models predicting building-emissions trajectories across 200+ building portfolios, turning them into interpretable LL97 compliance recommendations for 50+ clients.
    • Engineered pipelines to scrape, transform and aggregate multi-source energy data; built client-facing Power BI dashboards and supported LL97 compliance filings and payment processing for 900+ properties in coordination with DOB systems.
    • Automated client communications, document workflows and file management for 1,000+ clients with Python tools, cutting manual processing time by more than 50%.
    • Python
    • Statistical modeling
    • ETL
    • Power BI
    • Automation
  3. 2025

    Applied AI Intern

    at Aarohaa.AI

    May 2025 – Jun 2025
    New York, NY

    • Designed and deployed a JSON-parsing agent that normalizes LLM outputs across data formats, keeping inter-agent communication consistent in a multi-agent RAG / agentic pipeline.
    • Built FastAPI endpoints that connect AI agents to live dashboards for real-time processing, decision support and transparency in production.
    • Extracted, transformed and validated backend SQL datasets, engineering new columns and persistence logic for reliable storage across agent operations.
    • FastAPI
    • RAG
    • Agentic AI
    • LLMs
    • SQL
  4. 2023

    Campaign Strategist & Data Lead

    at DRUM: Desis Rising Up and Moving

    Apr 2024 – Jan 2026
    Bronx, NY

    Previously: Fellow, Sep – Dec 2023 · Chapter Committee Leader (part-time), Dec 2023 – Apr 2024

    • Designed and analyzed 100+ community surveys and maintained a 500+ member database, turning social-media, news and policy data into visualizations and reports for non-technical decision-makers.
    • Researched climate change and climate disasters across South Asia, and turned the analysis into data visualizations and teaching materials for community climate education, work I later extended into my climate research project.
    • Led multiple campaigns and designed leadership-development programs for youth and adult members; hosted workshops and trained members to lead their own sessions.
    • Managed members and volunteers through conflict and crisis, keeping campaigns on track under pressure.
    • Survey design
    • Data visualization
    • Climate research
    • Leadership development
    • Training & facilitation
  5. 2023

    Data Science Immersive Bootcamp Fellow

    at General Assembly

    Sep 2023 – Dec 2023
    New York, NY

    • 500+ hours of hands-on training in Python and SQL, exploratory data analysis, classical statistical modeling (study design, model evaluation, linear and logistic regression) and machine learning, from decision trees and random forests to NLP and neural networks.
    • Applied it end to end in three projects, each a written technical report plus a presentation built around a stakeholder question.
    • Built a climate and life-expectancy analysis of the US and Bangladesh against 190+ economies; an automated valuation model for residential property with conformal price intervals and out-of-time validation; and a Reddit NLP classifier (TF-IDF, logistic regression, Naive Bayes) that I later rebuilt into ModSieve, a moderation-triage tool.
    • Python
    • SQL
    • pandas
    • SciPy
    • scikit-learn
    • Statistical modeling
    • NLP
  • 200+building portfolios modeled for emissions at EN-POWER
  • 1,000+clients’ workflows automated with Python tools
  • 270+unit & E2E tests passing across my merged Kubeflow PRs
  • 190+economies reconciled from 6 public data sources, 1960–2024

03 projects

Things I’ve built.

open source · cncf

Kubeflow

Org member & contributor to the open-source ML platform, Jan 2026 – present.

  • Shipped merged PRs migrating Python clients off deprecated APIs to v1 per KEP-0004 across kubeflow/hub and kubeflow/sdk, regenerating OpenAPI clients and validating 270+ unit/E2E tests with pytest, mypy and ruff.
  • Designed and shipped a decoupled upload_artifact utility (Python, Pydantic) for S3 and OCI object storage, isolating artifact uploads from registry logic.
  • Building a Go client and backend proxy routing UI traffic to the v1 Model Registry, Model Catalog and MCP Catalog APIs; co-authored the SDK 0.4.0 release blog post.
  • Python
  • Go
  • Pydantic
  • OpenAPI
  • pytest

// applied machine learning · end-to-end, from data to result

Climate Risk & Emissions

US and Bangladesh vs. 190+ economies

Emissions drivers, backtested forecasts, carbon-budget scenarios to 2100 and income at risk from warming, reconciled across six public sources on 2024 data. Finds a 23× per-person emissions gap, and that even fast decarbonisation overshoots a 2 °C budget.

  • Python
  • statsmodels
  • scikit-learn
  • Plotly.js

Real-Estate Valuation & Risk

Ames, Iowa · 2,412 arm’s-length sales

An automated valuation model for residential property — the class of model a lender uses to price collateral. Every estimate carries an 80% price range whose coverage is measured, not assumed. LightGBM and regularised regression over a DuckDB feature layer, validated out-of-time: 5.5% median error, 77% of homes within ±10%. Pricing that valuation uncertainty into a loan book adds 85% to expected credit losses.

  • LightGBM
  • Conformal prediction
  • DuckDB SQL
  • Model risk

ModSieve

Reddit moderation triage · NLP · in progress

A first-pass moderation tool: it resolves the posts it’s confident about and sends the rest to a human moderator. Tuned to hold a fixed precision rather than chase accuracy, it auto-resolves 22.5% of posts at 99.2% precision, and scored 0.820 against my own blind labelling at 0.802. Labels come from real moderator removals, collected hourly across 8 communities.

  • NLP
  • scikit-learn
  • Decision thresholds
  • Trust & safety
View on GitHub

03 projects · flagship · in development

Sera: reading, engineered to scale.

“In an environment engineered for scrolling, sustained reading is hard.”

Sera is my attempt to make deep reading as frictionless as the things competing with it, and it’s where my neuroscience and computer science backgrounds meet. Upload a PDF and Sera is designed to read it aloud in an original, humanized voice while highlighting every word in sync, answer questions about whatever you highlight, and learn what you should read next.

I’m building it for its first 100 users and architecting it as though it will have 10,000, to see how larger infrastructure handles more data and more people, not just to ship the feature set.

  • 01

    Word-synced reading

    Per-word bounding boxes from the PDF are overlaid on the page, so the highlight follows the voice. Click any word and playback jumps there.

  • 02

    An original voice

    A synthetic character designed from scratch with StyleTTS 2: warm, calm, measured. No real person’s voice is cloned.

  • 03

    Ask Sera

    Highlight a passage and ask. A LangGraph agent retrieves context from the document with pgvector, then responds.

  • 04

    Reading intelligence

    Reading speed and dwell time shape recommendations from Semantic Scholar, arXiv and Open Library.

interactive

Drag to scale it.

Same product, four sizes. See what changes in the infrastructure as the user count grows.

100simulated users

BrowserReact + TS CloudFrontedge audio ALBround-robin FastAPI2 Fargate tasks Redis cacheTTS cache > 80% hit Postgres + pgvectorRDS · pooled Read replicarecs queries SQS queueTTS jobs TTS workers1 worker S3PDFs + audio

Staging · AWS ECS

p95 latency target
< 1 s API · < 3 s TTS
TTS cache hits
> 80%

What changes

    Design targets from the Sera architecture plan. Load tests at each tier will replace them with measured results, and I’ll publish the failures too.

    04 skills

    The toolkit.

    $ cat skills.txt

    // ai & machine learning

    • Generative AI
    • LLMs
    • Agentic frameworks
    • MCP
    • API integration
    • Graph neural networks
    • Statistical modeling
    • NLP
    • Neural networks
    • Predictive modeling
    • Time-series forecasting
    • Feature engineering

    // languages

    • Python
    • SQL
    • Go
    • TypeScript
    • C++
    • HTML / CSS

    // frameworks

    • PyTorch
    • PyTorch Geometric
    • TensorFlow
    • scikit-learn
    • LightGBM
    • statsmodels
    • LangChain
    • LangGraph
    • PySpark
    • FastAPI
    • Pydantic
    • React
    • Streamlit

    // data engineering & viz

    • PostgreSQL + pgvector
    • DuckDB
    • Redis
    • Web scraping
    • ETL pipelines
    • Multi-source data integration
    • Power BI
    • Matplotlib
    • Seaborn
    • Plotly
    • Interactive dashboards

    // cloud, devops & tools

    • AWS (ECS, S3, SQS)
    • Terraform
    • Docker
    • Git / GitHub
    • pytest
    • Supabase
    • Vercel
    • Claude Code
    • Google AI Studio

    $

    05 contact

    Let’s build
    something together.

    I’m open to roles in applied AI, machine learning, data and software engineering. The fastest way to reach me is email.