identity:
name: "Mohd Shami"
role: "AI / ML Engineer & Data Analyst"
location: "Moradabad, Uttar Pradesh, India"
education: "B.Tech in Data Science — Teerthanker Mahaveer University (Expected 2027)"
cgpa: "8.5 / 10 — Rank 1 in Academic Cohort"
focus: ["Machine Learning", "Deep Learning", "NLP", "MLOps", "Data Analytics"]
currently_exploring: ["Transformers", "Large Language Models", "Cloud ML Deployment"]
philosophy: "Data has a story — my job is to tell it well." |
I design and deploy end-to-end machine learning systems — from raw, messy datasets to production-ready models. My work spans predictive modeling, statistical analysis, and interactive BI dashboards, with hands-on experience across the XGBoost → Flask deployment pipeline and classical NLP/text-classification systems.
| Project | Description | Stack |
|---|---|---|
| Revive | ML-powered disease prediction web app. Improved diagnostic accuracy from 72% to 85% through advanced feature engineering and hyperparameter tuning. | Python Flask XGBoost Bootstrap |
| HamOrSpam Classifier | Email spam classifier using TF-IDF vectorization and Logistic Regression with SMOTE balancing, achieving over 98% accuracy with strong recall on imbalanced data. | Python Scikit-learn TF-IDF SMOTE |
| FarmAIQ | Smart agriculture system combining crop recommendation and plant disease detection pipelines, tuned with cross-validation for better generalization. | Python Random Forest CNN TensorFlow |
| ChurnShield AI | Customer churn prediction system built to flag at-risk customers ahead of time using classical ML classifiers. | Python Scikit-learn |
| CodeOrbit | A curated collection of 850+ DSA problems solved in Python, organized topic-wise for structured practice and revision. | Python DSA |
| Aug 2025 – Oct 2025 |
Data Science Intern — Codec Technologies India • Developed an XGBoost diagnostic model, improving accuracy from 72% to 85% • Ran end-to-end EDA, feature engineering, and data-cleaning pipelines on clinical datasets • Implemented reproducible training scripts with versioned experiments in Git XGBoost Python Git EDA
|
| Jun 2025 – Aug 2025 |
Python Developer Intern — CodTech IT Solutions Pvt. Ltd. • Built reusable Python ETL scripts (Pandas, NumPy) to clean heterogeneous data sources • Automated data validation and reporting templates, improving reproducibility • Reduced manual preprocessing time through automated pipelines Python Pandas NumPy ETL
|
|
B.Tech in Data Science Teerthanker Mahaveer University, Moradabad Expected 2027 · CGPA 8.5 / 10 · Rank 1 in Academic Cohort |
| Certification | Issuer | Year |
|---|---|---|
| SQL (Advanced) | HackerRank | 2026 |
| Getting Started with Data | IBM | — |
| AI – Machine Learning Engineer | Reliance Foundation | 2025 |
| Deep Learning Certificate | Simplilearn | 2025 |
| Problem Solving (Intermediate) | HackerRank | 2025 |
| Python (Basics) | HackerRank | 2025 |
| Metric | Value |
|---|---|
| Total Solved | 131 problems (122 Python3 · 8 Python · 1 MySQL) |
| Global Rank | 1,278,009 |
| Active Badge | 50 Days Badge 2026 |
| Advanced Topics | Dynamic Programming ×14 · Divide & Conquer ×6 · Backtracking ×3 |
| Intermediate Topics | Hash Table ×25 · Math ×23 · Greedy ×15 |
| Fundamental Topics | Array ×70 · String ×38 · Sorting ×19 |
| DSA Practice Archive | 850+ problems (see the CodeOrbit repository) |
Open to collaborations in Machine Learning, NLP, and Data Analytics — reach out for internships, open-source projects, or a good technical conversation.
If any of my work or repositories helped you, consider dropping a star or a coffee — it goes a long way for an independent learner.
Designed and maintained by Mohd Shami · Built with Markdown, SVG, and a lot of coffee · © 2026

