Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions Excel-CSV-File-Merger-main/LICENSE
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
MIT License

Copyright (c) 2026 Tuff-Tech-coder

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
157 changes: 157 additions & 0 deletions Excel-CSV-File-Merger-main/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,157 @@
# Excel & CSV File Merger

![Python](https://img.shields.io/badge/Python-3.11+-3776AB?logo=python&logoColor=white)
![pandas](https://img.shields.io/badge/pandas-3.x-150458?logo=pandas&logoColor=white)
![openpyxl](https://img.shields.io/badge/openpyxl-3.x-1D6F42)
![Tests](https://img.shields.io/badge/tests-34%20passing-brightgreen)
![License](https://img.shields.io/badge/License-MIT-green)

A data-normalization pipeline that scans a folder for Excel and CSV files and intelligently merges them into a single, clean master workbook. Built specifically for the **messy reality** of business data: mismatched column orders, inconsistent naming, currency stored as text, stray whitespace, blank rows, and duplicates.

---

## Why it's useful

Combining spreadsheets from different teams by hand is slow and error-prone — column names never match, totals are formatted as `"$5,400.00"`, and duplicates creep in. This tool encodes those fixes once and applies them consistently, turning a pile of inconsistent files into one analysis-ready dataset with an audit trail.

---

## Features

- **Auto-discovery** of every `.xlsx`, `.xls`, and `.csv` file in a folder.
- **Column-alias mapping** — a configurable dictionary unifies variants (`"Sales Amount"`, `"Amount"`, `"Salary"` → `revenue`; `"Territory"` → `region`) so files merge on meaning, not exact spelling.
- **Currency normalization** — `"$5,400.00"` becomes `5400.0` for real math.
- **Cleanup pass** — strips cell whitespace and drops entirely empty rows.
- **Graceful column alignment** — files missing columns still merge cleanly, with gaps filled rather than erroring.
- **Configurable deduplication** of identical rows (`--no-dedup` to disable).
- **Source tracking** — a `source_file` column records each row's origin.
- **Polished two-sheet output** — a formatted *Merged Data* sheet (styled header, alternating row shading, auto-fit columns) plus a *Merge Summary* sheet with per-file counts, duplicates removed, and final totals.

---

## Tech stack

`Python` · `pandas` · `openpyxl` · `xlrd` · `argparse` · `logging` · `regex`

---

## Project structure

```
├── excel_merger.py # Main script
├── create_samples.py # Generates the demo input files
├── requirements.txt # Python dependencies
├── tests/ # 26 unit tests
├── merged_master.xlsx # Generated: merged output (gitignored)
├── merger.log # Generated: application log
└── sample_input/
├── sales_q1.xlsx # Standard sales data
├── sales_q2.xlsx # Different column order + extra column + a duplicate
├── hr_employees.csv # Missing columns, different naming
├── marketing_leads.csv # Currency-as-text, whitespace, varied date formats
└── ops_data.xlsx # Renamed columns + blank rows
```

---

## Setup

```bash
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
```

---

## Usage

```bash
# Merge the included sample files
python excel_merger.py --input ./sample_input --output merged_master.xlsx

# Merge your own folder, keep duplicates
python excel_merger.py --input /path/to/folder --no-dedup
```

**Options:** `--input` · `--output` · `--no-dedup`

---

## How it works

1. Discovers all supported files in the input folder.
2. Loads each, normalizes column names through the alias map, cleans cells, normalizes currency, drops empty rows, and tags rows with their source file.
3. Concatenates everything (pandas aligns on column name, filling gaps).
4. Optionally deduplicates on all columns except the source tag.
5. Writes a formatted workbook with both the merged data and a summary sheet.

---

## Customizing: the `COLUMN_ALIASES` extension point

This is the most reusable idea in the repo. `COLUMN_ALIASES` is a single
dictionary mapping every known source header onto a canonical name:

```python
COLUMN_ALIASES: dict[str, str] = {
"full name": "name",
"employee name": "name",
"sales amount": "revenue",
"amount": "revenue",
"salary": "revenue",
"territory": "region",
...
}
```

Teaching the merger a new file format is **one line in this dict** — no other
code changes. Everything downstream (currency normalization, deduplication,
column alignment, the summary sheet) keys off canonical names, so the entire
pipeline picks up the new variant automatically. Schema drift becomes
configuration rather than a code change.

### Two caveats worth knowing

**Aliasing can collide.** Several headers intentionally map to the same
canonical name — `Amount`, `Total` and `Salary` all become `revenue`. If a
*single file* contains two of them, they cannot both be `revenue`. `dedupe_columns()`
renames the second to `revenue_2` and logs a warning naming the file:

```
marketing_leads.csv: duplicate canonical column 'revenue' renamed to 'revenue_2'.
Review COLUMN_ALIASES if these should be merged.
```

Nothing is silently dropped, and nothing crashes — you get both columns plus a
prompt to decide whether that mapping was right for your data. Note that only
the canonical `revenue` column gets currency normalization; `revenue_2` is left
as-is precisely because the tool should not guess which one you meant.

**`"first name" → "name"` is lossy.** If a file has separate `First Name` and
`Last Name` columns, only the first is mapped to `name` and the surname stays
under its own column. That is a deliberate simplification, not a bug — but if
your data splits names that way, either remove that alias or pre-join the two
columns before merging.

---

## Development

```bash
pip install -r requirements.txt
pip install pytest ruff

pytest -q # 34 tests
ruff check .
```

The suite covers alias normalization, currency parsing, whitespace cleaning,
deduplication and source-file tagging, including a regression test for the
alias-collision crash described above.

---

## Possible extensions

Add per-column type coercion, configurable merge keys for true joins (not just concatenation), a dry-run preview, or output to Parquet/SQL for larger datasets.
Binary file not shown.
Binary file not shown.
133 changes: 133 additions & 0 deletions Excel-CSV-File-Merger-main/create_samples.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,133 @@
"""Create five deliberately messy sample input files for the Excel Merger demo."""

import csv
from pathlib import Path

import pandas as pd

OUT = Path(__file__).parent / "sample_input"


def create_samples(output_folder: Path = OUT) -> None:
"""Generate the demo workbooks and CSV files in ``output_folder``."""
output_folder.mkdir(parents=True, exist_ok=True)

# ── File 1: sales_q1.xlsx ─────────────────────────────────────────────
# Standard sales data, some mixed-case headers
df1 = pd.DataFrame({
"Name": ["Alice Johnson", "Bob Smith", "Carol White", "Dave Brown", "Eve Davis"],
"Email": [
"alice@corp.com",
"bob@corp.com",
"carol@corp.com",
"dave@corp.com",
"eve@corp.com",
],
"Revenue": [12500.00, 8750.50, 23400.00, 5600.75, 19800.00],
"Region": ["North", "South", "North", "West", "East"],
"Date": ["2024-01-15", "2024-01-20", "2024-02-03", "2024-02-14", "2024-03-01"],
"Department": ["Sales", "Sales", "Sales", "Sales", "Sales"],
})
df1.to_excel(output_folder / "sales_q1.xlsx", index=False)

# ── File 2: sales_q2.xlsx ─────────────────────────────────────────────
# Same data type but different column order and some extra columns
df2 = pd.DataFrame({
"email": [
"frank@corp.com",
"grace@corp.com",
"henry@corp.com",
"irene@corp.com",
"jake@corp.com",
"alice@corp.com",
],
"REVENUE": [7200.00, 15300.00, 9800.00, 22100.00, 6400.00, 12500.00],
"Region": ["South", "East", "North", "West", "South", "North"],
"name": ["Frank Lee", "Grace Kim", "Henry Chen", "Irene Park", "Jake Wu", "Alice Johnson"],
"Date": [
"2024-04-10",
"2024-04-22",
"2024-05-08",
"2024-05-19",
"2024-06-01",
"2024-01-15",
],
"Notes": ["Top performer", "", "Needs review", "Excellent", "", "Duplicate entry"],
"Department": ["Sales", "Sales", "Sales", "Sales", "Sales", "Sales"],
})
df2.to_excel(output_folder / "sales_q2.xlsx", index=False)

# ── File 3: hr_employees.csv ─────────────────────────────────────────
# CSV format, missing Revenue and Date columns, extra HR-specific columns
df3 = pd.DataFrame({
"Name": [
"Laura Moss",
"Mike Stone",
"Nancy Hill",
"Oscar Reed",
"Paula Marsh",
"Quinn Ford",
"Rachel Moore",
],
"Email": [
"laura@corp.com",
"mike@corp.com",
"nancy@corp.com",
"oscar@corp.com",
"paula@corp.com",
"quinn@corp.com",
"rachel@corp.com",
],
"Department": ["HR", "HR", "Engineering", "Engineering", "Marketing", "Marketing", "HR"],
"Region": ["East", "West", "North", "South", "East", "West", "North"],
"Hire Date": [
"2021-03-15",
"2019-07-01",
"2022-11-20",
"2020-04-05",
"2023-01-10",
"2018-09-30",
"2024-02-15",
],
"Salary": [55000, 72000, 98000, 115000, 61000, 68000, 52000],
})
df3.to_csv(output_folder / "hr_employees.csv", index=False)

# ── File 4: marketing_leads.csv ──────────────────────────────────────
# CSV with inconsistent formatting: extra whitespace, mixed number formats
rows = [
[" Name ", "Email", "Revenue", " Region ", "Date", "Department"],
["Sam Turner ", "sam@leads.com", "$5,400.00", " North ", "2024/01/08", "Marketing"],
[" Tina Brooks", "tina@leads.com", "8200", "East", "2024-02-14", "Marketing"],
["Uma Patel ", "uma@leads.com", "$11,750.50", "West ", "March 5, 2024", "Marketing"],
["Vince Hall", "vince@leads.com", "3100.0", "South", "2024-04-22", "Marketing"],
[" Wendy Fox", "wendy@leads.com", "$19,900", "North", "2024-05-30", "Marketing"],
["Xavier Long", "xavier@leads.com", "7650.25", " East", "2024-06-15", "Marketing"],
]
with (output_folder / "marketing_leads.csv").open("w", encoding="utf-8", newline="") as f:
writer = csv.writer(f)
writer.writerows(rows)

# ── File 5: ops_data.xlsx ────────────────────────────────────────────
# Different column names, some blank rows, purely numeric revenue
df5 = pd.DataFrame({
"Full Name": ["Yara Singh", "Zach Adams", None, "Amy Clarke", "Brian Duke"],
"Contact Email": [
"yara@ops.com",
"zach@ops.com",
None,
"amy@ops.com",
"brian@ops.com",
],
"Sales Amount": [14200.00, 9900.50, None, 31000.00, 7800.25],
"Territory": ["West", "North", None, "South", "East"],
"Transaction Date": ["2024-03-18", "2024-04-02", None, "2024-05-11", "2024-06-29"],
"Team": ["Operations", "Operations", None, "Operations", "Operations"],
})
df5.to_excel(output_folder / "ops_data.xlsx", index=False)

print(f"[OK] Created 5 sample input files in {output_folder}")


if __name__ == "__main__":
create_samples()
Loading
Loading