Hey DataGeeks, I'm Pavan π
π Location : Austin, TX, USA
π Mobile : +1 (929) 278-4589
βοΈ Email : pavan.yellathakota.ds@gmail.com
Linkedin : https://linkedin.com/in/yellatp
GitHub : https://github.com/yellatp
Website : onlynerds.win
AI/ML Engineer with 4+ years of experience across the full arc of the data career path. I began in customer and consumer analytics as a Data Analyst, grew into a Data Scientist role building supply chain analytics, inventory management systems, and seasonal and event-based forecasting models, and now work as an AI/ML Engineer developing and deploying production ML and Generative AI systems. I have consulted for 10+ local businesses through Clarkson University, co-founded my own venture, Alphonso AI, and launched onlynerds.win, a resources and job search platform for job seekers. I currently work as an AI/ML Engineer at USK Systems, consulting and building products for Fortune 100 and Fortune 500 clients as well as promising AI startups.
| Domain | Stack |
|ββββββββββ-|ββ-|
| Languages & Databases |
|
| AWS Cloud Data |
|
| ML Frameworks |
|
| Tools & Visualization |
|
Remote | 2026 β Present
Potsdam, NY | Jul 2025 β Present
Key Technologies Used
Potsdam, NY | Sep 2024 β Apr 2025
Key Technologies Used
Potsdam, NY | Oct 2023 β May 2025
Key Technologies Used
Hyderabad, India | Jul 2022 β Dec 2022
Key Technologies Used
Pune, India | Sep 2021 β May 2022
Key Technologies Used
Alphonso AI β Co-Founder and Founding ML Engineer. I helped start Alphonso AI, backed by the Shipley Center for Innovation, and lead the ML and backend engineering side of the company, covering everything from the 0β1 backend architecture to the search and recommendation systems described in the experience section above. Full details on this work are in Professional Experience.
onlynerds.win β Only Nerds, an ATS Intelligence Platform I conceived, designed, and built on my own, outside of any employer or client engagement. It maps 115,008+ unique company records to the applicant tracking system (ATS) each one runs on, including Greenhouse, Ashby, Lever, BambooHR, iCIMS, Workable, join.com, Personio, and 20+ more, so job seekers can identify a companyβs ATS and reach its job board in one click. The site runs on Astro and TypeScript, and I handle everything end to end: product direction, data pipeline, frontend build, and deployment.
Each project below reflects the actual scope of the underlying repository, including the problem it addresses, the sector it applies to, and the technical approach behind it.
Detoxify Telugu
A fine-tuned BERT-based language model for hate speech detection in Telugu and Tenglish (Telugu written in Latin script). Large general-purpose language models are trained overwhelmingly on high-resource languages and routinely miss dialect-specific slang and mixed-script text, which lets harmful content slip past moderation in regional markets. This project fine-tunes a transformer model specifically on Telugu and Tenglish text to close that gap.
BingeMax Recommendation Engine
An AI-powered movie recommendation system that combines content-based filtering, collaborative filtering, and cosine similarity scoring to generate personalized suggestions, served through a Streamlit interface backed by a FastAPI service layer.
Fintech Sales GAP Analysis
A Python-based analytics project built for fintech sales teams to quantify performance gaps across sales representatives, regions, and product lines, translating the findings into concrete, data-driven training recommendations for underperforming segments.
KonnectR
A full-stack web application that connects academia and industry, letting students, professors, and professionals collaborate in one place. It includes an asynchronous chat system, listings for jobs and research opportunities, and role-based user management, and was later extended into a second build, KonnectR_flask_fullstack_app, adding analytics and posting modules on top of the same concept.
CUDA vs CPU Showdown
A hardware-aware performance benchmarking study comparing GPU-accelerated data processing against traditional CPU workflows, using RAPIDS (cuDF) and DuckDB across datasets of over 1M rows on an NVIDIA GTX 1650. The results document a 3x to 10x throughput improvement from GPU acceleration, making the case for hardware-aware pipeline design in data engineering work.
XIFTY
An edge-native, desktop-first intelligence suite built to scout, analyze, and manage content creators at scale, designed for fast local performance rather than a purely cloud-dependent architecture.
AtmosDB β A unified SDK that brings Cloudflare D1, Vectorize, and R2 together behind one interface, aimed at building global, AI-native backends at the edge without the usual dashboard overhead. (TypeScript)
Customer Acquisition Cost Analysis β An in-depth analysis of customer acquisition cost across marketing channels for 2023, evaluating CAC, conversion efficiency, and spend effectiveness to guide marketing budget decisions. (Python, Pandas)
Text Analysis using NLP and LDA β An NLP project covering sentiment analysis, named entity recognition, and topic modeling with LDA, applied to unstructured text corpora. (Python, NLTK, Gensim)
Fake News Classifier β A machine learning system that classifies news articles as fake or real using Naive Bayes and SVM models, with a full text preprocessing and evaluation pipeline. (Python, Scikit-learn)
GenZ Career Preferences Report β A data-driven study of career preferences among Gen Z professionals, based on survey responses from 230+ participants, presented through an interactive dashboard. (Python, Pandas, Plotly)
Content Strategy Analysis: Netflix β An analysis of Netflixβs 2023 viewership trends, examining content performance, audience preferences, and seasonal patterns to inform content strategy decisions. (Python, Pandas, Matplotlib, Seaborn)
Supply Chain Analysis β An in-depth look at supply chain data covering product performance, inventory, suppliers, logistics, and customer behavior, with interactive visualizations. (Python, Pandas, NumPy, Plotly)
PreOwned Cars Price Prediction β A regression project predicting used car prices across US regions using Multi-Linear Regression, Decision Trees, Random Forests, and XGBoost, with K-Means clustering for regional segmentation. A follow-up version, V2.0, adds a manufacturer-depreciation-to-odometer metric to improve prediction accuracy. (Python, Scikit-learn, XGBoost)
Synthetic Data Generator β A configurable tool with a Streamlit interface for generating synthetic datasets, useful for testing pipelines and augmenting training data without relying on real user data. (Python, Streamlit)
Website A/B Testing β A data-driven comparison of website performance under light and dark themes, using statistical tests, correlation analysis, and predictive modeling to identify which theme drives stronger engagement. (Python, Pandas, Seaborn, Scikit-learn)
Last Updated: 2026 by PAVAN YELLATHAKOTA </sub> </p>