Skip to content
View RISHIKKASULA's full-sized avatar

Block or report RISHIKKASULA

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
RISHIKKASULA/README.md

Rishik Kasula — Machine Learning & Data Engineer

I like systems that can prove they work: calibrated models instead of raw scores, LLM output checked against ground truth, pipelines that quarantine bad data instead of hiding it.

M.S. Computer Science, Data Science concentration, UNC Charlotte — May 2026.

Shipped

Three public repos. Sole author. Each one ships with tests, CI, and a tagged release — or it doesn't ship.

fraudscore — v1.0.0

Card-fraud scoring that prices every decision: review a transaction when P(fraud) × amount clears the cost of reviewing it. No fitted threshold, just the expected-cost rule.

Under leak-free chronological evaluation the amount-aware rule saved $84.87 per 10,000 transactions, 95% CI [$20.08, $187.47] — about 13% cheaper than the best single-threshold baseline. It reviews fewer frauds than that baseline and still wins, because it wins on dollars rather than on counts. One finding worth the click: a plain logistic model beat gradient boosting once the evaluation stopped leaking.

Python · scikit-learn · FastAPI · DuckDB · Docker · bootstrap CIs

filinglens — v0.1.0

An evaluation harness, not an extractor. Every US filer publishes its numbers twice — once in prose, once in structured XBRL — so the company's own filing grades the model automatically. No human answer key.

On headline-figure extraction with the financial statements already in the model's context, llama3.1-8B reached 87.8%, 95% CI [81.5, 93.8], across ten large-cap clean us-gaap annual filers. That's the easy end of the task, and the framing travels with the number. A 600-call deterministic grid, 588 scored, all 230 failures hand-labeled into a six-label taxonomy.

Python · XBRL / EDGAR · BM25 · cluster bootstrap · mypy strict

tickflow — v0.9

A quality gate for Kafka-compatible streams, measured by what it prevents. Declarative per-topic contracts enforced inline; violations quarantined with self-describing envelopes rather than dropped.

Same faulted fixture replayed through the same bar builder — gates off, 1,076 violated bars of 15,061 got through; gates on, zero, with the output bit-identical to the ground-truth valid projection. Every figure is fixture-scale: a seeded synthetic fixture with a chosen fault mix, never live traffic. 194 tests, 99.4% coverage on core modules, mypy strict.

A stated v0.9, not a 1.0 — the quarantine/replay CLI got cut under a slip valve frozen before the deadline, and the ADR saying so is in the repo. Shipping the roadmap honestly is the point.

Python 3.12 · Kafka / Redpanda · Avro · DuckDB · GitHub Actions

Background

Graduate research assistant, Jan–May 2026. Structural analysis of code refactorings — mutual-information-derived scoring and influence clustering over code property graphs. About 3,500 lines of git-verified contributions. The repo is private faculty research, so the three above are the ones you can actually read.

Instructional assistant, Sep–Dec 2025. Graduate course on software engineering for AI-enabled systems. Office hours, project mentoring across the ML lifecycle, grading.

Earlier. Real-time Apache Kafka + ZooKeeper ingestion and Power BI KPI dashboards in a banking-IT setting, plus a local open-source LLM deployment.

Tools

Python · SQL · scikit-learn · pandas · Kafka · Avro · DuckDB · FastAPI · Docker · GitHub Actions · pytest · mypy · Power BI · AWS Glue / S3 · Git


Open to full-time ML Engineer and Data Engineer roles in the U.S.

Portfolio · LinkedIn · contact.rishikkasula@gmail.com

Pinned Loading

  1. filinglens filinglens Public

    XBRL-graded evaluation harness measuring how reliably a local LLM extracts financial figures from SEC filings.

    Python

  2. fraudscore fraudscore Public

    Calibrated card-fraud scoring with expected-cost decisions — evaluated in dollars with bootstrap CIs, served over FastAPI

    Python

  3. tickflow tickflow Public

    Python