Research Assistant in Data Science & Educational AI at UNC School of Education, building open-source AI tutoring infrastructure for mathematics education.
Background
MS in Data Science Engineering (GPA 3.7) from the University of Houston, with a B.Tech in AI & Machine Learning.
I build across the full AI stack — from training deep learning models and designing multi-agent CrewAI pipelines to deploying production Streamlit apps backed by vector databases.
Currently a Research Assistant in Data Science & Educational AI at UNC School of Education, building reusable, open-source AI tutoring infrastructure for mathematics education.
Beyond AI, I ship full-stack systems end-to-end — a Next.js/TypeScript/GraphQL/Kubernetes app (FlowBoard) and a React/Node/PostgreSQL platform (DevLink) — to stay sharp across the whole stack, not just the model layer.
Where I've Worked
- Supporting development of reusable, open-source AI tutoring infrastructure for mathematics education.
- Working across data preparation, analysis, model evaluation, software development, and documentation.
- Collaborating with an interdisciplinary team spanning data science, AI, learning sciences, and education.
- Extending prior PhET Learning Progression feedback pipeline work into a broader open-source tutoring platform.
- Built an AI chatbot integrated with PhET physics simulations to analyze 100+ student interaction responses.
- Designed Python pipelines to classify responses via Learning Progression rubrics — improving accuracy by 20%.
- Implemented LLM evaluation pipelines with self-consistency voting, boosting reliability by 25%.
- Automated personalized feedback generation, cutting manual grading effort by 40%.
- Built scalable processing workflows, improving pipeline efficiency by 30%.
- Built ML models using Scikit-learn and Pandas to analyze 5,000+ anonymized healthcare records.
- Performed feature engineering on 50+ clinical features, improving model performance by 15%.
- Trained Random Forest and Gradient Boosting models achieving 85%+ prediction accuracy.
- Created data visualization dashboards, improving reporting efficiency by 30%.
- Optimized preprocessing pipelines, reducing data processing time by 25%.
Academic Background
What I Work With
Projects
- 3 CrewAI agents cut report interpretation time by 70%
- LangChain + ChromaDB for 50+ patient records & trend analysis
- Streamlit UI with REST APIs for real-time insights
- LLM evaluation pipelines improved reliability by 25%
- Automated feedback reduced evaluation time by 40%
- Scalable Python modules for large-batch processing
- Preprocessed and augmented 1,000+ retinal images
- Improved generalization by 15% via hyperparameter tuning
- 88% accuracy across all retinopathy severity stages
- Apollo GraphQL API layer over a Next.js/TypeScript frontend
- Containerized and deployed via Kubernetes
- CI/CD pipeline automated with GitHub Actions
- React frontend on Vercel, Node/Express backend on Render
- PostgreSQL (Neon) + Redis (Upstash) for persistence and caching
- GPT-4o-mini API integration for smart features