yellatp

Pavan Yellathakota Hey DataGeeks, I'm Pavan πŸ‘‹
I'm an AI/ML Engineer who understands the data world from the ground up. I started in customer and consumer analytics as a Data Analyst, moved into supply chain analytics, inventory management, and seasonal and event-based demand forecasting as a Data Scientist, and now build and deploy production ML models and Generative AI systems. Along the way I consulted for 10+ local businesses through Clarkson University, including HAVK Mladost, co-founded my own venture, **Alphonso AI** (alphonso.app), and launched **onlynerds.win**, a resources and job search platform for job seekers. I'm currently an AI/ML Engineer at **USK Systems**, consulting and building products for Fortune 100 and Fortune 500 clients as well as promising AI startups.

AI/ML Engineer | Data Scientist | Generative AI

πŸ“ Location : Austin, TX, USA
πŸ“ž Mobile : +1 (929) 278-4589
βœ‰οΈ Email : pavan.yellathakota.ds@gmail.com
Linkedin : https://linkedin.com/in/yellatp
GitHub :   https://github.com/yellatp
Website :   onlynerds.win



πŸ‘¨β€πŸ’» Professional Summary

AI/ML Engineer with 4+ years of experience across the full arc of the data career path. I began in customer and consumer analytics as a Data Analyst, grew into a Data Scientist role building supply chain analytics, inventory management systems, and seasonal and event-based forecasting models, and now work as an AI/ML Engineer developing and deploying production ML and Generative AI systems. I have consulted for 10+ local businesses through Clarkson University, co-founded my own venture, Alphonso AI, and launched onlynerds.win, a resources and job search platform for job seekers. I currently work as an AI/ML Engineer at USK Systems, consulting and building products for Fortune 100 and Fortune 500 clients as well as promising AI startups.


πŸ› οΈ Technical Skills

| Domain | Stack | |β€”β€”β€”β€”β€”β€”β€”β€”β€”β€”-|β€”β€”-| | Languages & Databases | Python Pandas NumPy Scikit-learn SQL PySpark PostgreSQL R TypeScript | | AWS Cloud Data | AWS S3 Athena Glue SageMaker Lambda Redshift | | ML Frameworks | XGBoost Hugging Face OpenAI NLTK Spacy | | Tools & Visualization | Tableau QuickSight Power BI Excel Git |


πŸ’Ό Professional Experience

USK Systems | AI/ML Engineer, Consultant

Remote | 2026 – Present

Alphonso AI, backed by Shipley Center for Innovation | Co-Founder & Founding ML Engineer

Potsdam, NY | Jul 2025 – Present

Key Technologies Used
Python FastAPI PostgreSQL Docker pgvector HuggingFace Gemini Vertex AI

Student Managed Investment Fund, Clarkson University | Graduate Quantitative Researcher

Potsdam, NY | Sep 2024 – Apr 2025

Key Technologies Used
Python BERT HuggingFace Vertex AI Pandas

HAVK Mladost (Elite Athletics Club) | Graduate Data Science Consultant

Potsdam, NY | Oct 2023 – May 2025

Key Technologies Used
AWS S3 Glue PySpark FastAPI Python

eAppSys Limited | Business Data Analyst

Hyderabad, India | Jul 2022 – Dec 2022

Key Technologies Used
Python Prophet SARIMAX Oracle OCI

Kantar GDC India | Data Analyst

Pune, India | Sep 2021 – May 2022

Key Technologies Used
Python PySpark Pandas SQL


πŸš€ Independent Ventures

Alphonso AI β€” Co-Founder and Founding ML Engineer. I helped start Alphonso AI, backed by the Shipley Center for Innovation, and lead the ML and backend engineering side of the company, covering everything from the 0β†’1 backend architecture to the search and recommendation systems described in the experience section above. Full details on this work are in Professional Experience.

onlynerds.win β€” Only Nerds, an ATS Intelligence Platform I conceived, designed, and built on my own, outside of any employer or client engagement. It maps 115,008+ unique company records to the applicant tracking system (ATS) each one runs on, including Greenhouse, Ashby, Lever, BambooHR, iCIMS, Workable, join.com, Personio, and 20+ more, so job seekers can identify a company’s ATS and reach its job board in one click. The site runs on Astro and TypeScript, and I handle everything end to end: product direction, data pipeline, frontend build, and deployment.


πŸ—οΈ Notable Projects

Each project below reflects the actual scope of the underlying repository, including the problem it addresses, the sector it applies to, and the technical approach behind it.

Flagship Projects

Detoxify Telugu
A fine-tuned BERT-based language model for hate speech detection in Telugu and Tenglish (Telugu written in Latin script). Large general-purpose language models are trained overwhelmingly on high-resource languages and routinely miss dialect-specific slang and mixed-script text, which lets harmful content slip past moderation in regional markets. This project fine-tunes a transformer model specifically on Telugu and Tenglish text to close that gap.

BingeMax Recommendation Engine
An AI-powered movie recommendation system that combines content-based filtering, collaborative filtering, and cosine similarity scoring to generate personalized suggestions, served through a Streamlit interface backed by a FastAPI service layer.

Fintech Sales GAP Analysis
A Python-based analytics project built for fintech sales teams to quantify performance gaps across sales representatives, regions, and product lines, translating the findings into concrete, data-driven training recommendations for underperforming segments.

KonnectR
A full-stack web application that connects academia and industry, letting students, professors, and professionals collaborate in one place. It includes an asynchronous chat system, listings for jobs and research opportunities, and role-based user management, and was later extended into a second build, KonnectR_flask_fullstack_app, adding analytics and posting modules on top of the same concept.

CUDA vs CPU Showdown
A hardware-aware performance benchmarking study comparing GPU-accelerated data processing against traditional CPU workflows, using RAPIDS (cuDF) and DuckDB across datasets of over 1M rows on an NVIDIA GTX 1650. The results document a 3x to 10x throughput improvement from GPU acceleration, making the case for hardware-aware pipeline design in data engineering work.

XIFTY
An edge-native, desktop-first intelligence suite built to scout, analyze, and manage content creators at scale, designed for fast local performance rather than a purely cloud-dependent architecture.

Additional Projects

AtmosDB β€” A unified SDK that brings Cloudflare D1, Vectorize, and R2 together behind one interface, aimed at building global, AI-native backends at the edge without the usual dashboard overhead. (TypeScript)

Customer Acquisition Cost Analysis β€” An in-depth analysis of customer acquisition cost across marketing channels for 2023, evaluating CAC, conversion efficiency, and spend effectiveness to guide marketing budget decisions. (Python, Pandas)

Text Analysis using NLP and LDA β€” An NLP project covering sentiment analysis, named entity recognition, and topic modeling with LDA, applied to unstructured text corpora. (Python, NLTK, Gensim)

Fake News Classifier β€” A machine learning system that classifies news articles as fake or real using Naive Bayes and SVM models, with a full text preprocessing and evaluation pipeline. (Python, Scikit-learn)

GenZ Career Preferences Report β€” A data-driven study of career preferences among Gen Z professionals, based on survey responses from 230+ participants, presented through an interactive dashboard. (Python, Pandas, Plotly)

Content Strategy Analysis: Netflix β€” An analysis of Netflix’s 2023 viewership trends, examining content performance, audience preferences, and seasonal patterns to inform content strategy decisions. (Python, Pandas, Matplotlib, Seaborn)

Supply Chain Analysis β€” An in-depth look at supply chain data covering product performance, inventory, suppliers, logistics, and customer behavior, with interactive visualizations. (Python, Pandas, NumPy, Plotly)

PreOwned Cars Price Prediction β€” A regression project predicting used car prices across US regions using Multi-Linear Regression, Decision Trees, Random Forests, and XGBoost, with K-Means clustering for regional segmentation. A follow-up version, V2.0, adds a manufacturer-depreciation-to-odometer metric to improve prediction accuracy. (Python, Scikit-learn, XGBoost)

Synthetic Data Generator β€” A configurable tool with a Streamlit interface for generating synthetic datasets, useful for testing pipelines and augmenting training data without relying on real user data. (Python, Streamlit)

Website A/B Testing β€” A data-driven comparison of website performance under light and dark themes, using statistical tests, correlation analysis, and predictive modeling to identify which theme drives stronger engagement. (Python, Pandas, Seaborn, Scikit-learn)



Last Updated: 2026 by PAVAN YELLATHAKOTA </sub> </p>