Spaces:
Runtime error
Runtime error
Commit ·
23e9906
1
Parent(s): 0d541e6
Add final files and prepare for deployment
Browse files- # .gitignore +17 -0
- README.md +74 -57
- __pycache__/api.cpython-312.pyc +0 -0
- __pycache__/models.cpython-312.pyc +0 -0
- api.py +29 -24
- app.py +27 -0
- generate_output.py +100 -0
- models.py +1 -2
- requirements.txt +16 -16
# .gitignore
ADDED
|
@@ -0,0 +1,17 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# .gitignore
|
| 2 |
+
.venv/
|
| 3 |
+
__pycache__/
|
| 4 |
+
*.pyc
|
| 5 |
+
*.pyo
|
| 6 |
+
*.pyd
|
| 7 |
+
.DS_Store
|
| 8 |
+
*.log
|
| 9 |
+
*.tmp
|
| 10 |
+
*.swp
|
| 11 |
+
|
| 12 |
+
# Optional: If dataset is too large or shouldn't be committed
|
| 13 |
+
# combined_emails_with_natural_pii.csv
|
| 14 |
+
# api_output_results.jsonl
|
| 15 |
+
|
| 16 |
+
# Optional: If model is very large and managed by Git LFS or stored elsewhere
|
| 17 |
+
# saved_models/
|
README.md
CHANGED
|
@@ -1,103 +1,120 @@
|
|
| 1 |
# Email Classification and PII Masking API
|
| 2 |
|
| 3 |
-
This project implements an API
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
|
| 5 |
## Project Structure
|
| 6 |
|
| 7 |
```
|
| 8 |
-
|
| 9 |
-
├──
|
| 10 |
-
|
| 11 |
-
├──
|
| 12 |
-
├──
|
| 13 |
-
├──
|
| 14 |
-
├──
|
| 15 |
-
├── .
|
| 16 |
-
├── .
|
| 17 |
-
|
|
|
|
|
|
|
| 18 |
```
|
| 19 |
|
| 20 |
## Setup
|
| 21 |
|
| 22 |
-
1. **Clone the repository:**
|
| 23 |
```bash
|
| 24 |
git clone <your-repo-url>
|
| 25 |
-
cd
|
| 26 |
```
|
| 27 |
-
|
| 28 |
-
2. **Create and activate a virtual environment:** (Recommended)
|
| 29 |
```bash
|
| 30 |
-
python -m venv venv
|
| 31 |
-
source venv/bin/activate
|
|
|
|
| 32 |
```
|
| 33 |
-
|
| 34 |
3. **Install dependencies:**
|
| 35 |
```bash
|
| 36 |
pip install -r requirements.txt
|
| 37 |
```
|
| 38 |
-
|
| 39 |
-
4. **Download NLP models:**
|
| 40 |
```bash
|
| 41 |
python -m spacy download en_core_web_sm
|
| 42 |
-
# Run inside python interpreter if needed:
|
| 43 |
-
# import nltk; nltk.download('punkt'); nltk.download('stopwords')
|
| 44 |
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
|
| 46 |
-
|
| 47 |
-
* You need to train the classification model first. Add your training data to the `data/` directory.
|
| 48 |
-
* Adapt and run the training logic (e.g., the `train_and_save_model` function in `models.py`, potentially moving it to a separate `train.py` script). This will create the necessary files in `saved_models/`.
|
| 49 |
-
* Example (modify as needed): `python models.py` (if you add training execution there) or `python train.py`
|
| 50 |
-
|
| 51 |
-
## Running the API Locally
|
| 52 |
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
|
| 56 |
-
|
| 57 |
-
|
| 58 |
-
|
| 59 |
-
The API will be available at `http://
|
| 60 |
-
|
| 61 |
-
## API Usage
|
| 62 |
|
| 63 |
-
|
| 64 |
|
| 65 |
-
|
| 66 |
|
| 67 |
-
**
|
| 68 |
|
| 69 |
-
```
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
|
|
|
|
|
|
|
| 73 |
```
|
| 74 |
|
| 75 |
-
**
|
| 76 |
|
| 77 |
```json
|
| 78 |
{
|
| 79 |
-
"input_email_body": "
|
| 80 |
"list_of_masked_entities": [
|
| 81 |
{
|
| 82 |
-
"position": [
|
| 83 |
"classification": "full_name",
|
| 84 |
-
"entity": "
|
| 85 |
},
|
| 86 |
{
|
| 87 |
-
"position": [
|
| 88 |
-
"classification": "
|
| 89 |
-
"entity": "
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 90 |
}
|
| 91 |
],
|
| 92 |
-
"masked_email": "
|
| 93 |
-
"category_of_the_email": "
|
| 94 |
}
|
| 95 |
```
|
| 96 |
|
| 97 |
-
##
|
| 98 |
|
| 99 |
-
|
|
|
|
|
|
|
| 100 |
|
| 101 |
-
##
|
| 102 |
|
| 103 |
-
(
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
# Email Classification and PII Masking API
|
| 2 |
|
| 3 |
+
This project implements an API service for classifying support emails into predefined categories while masking Personally Identifiable Information (PII) before classification.
|
| 4 |
+
|
| 5 |
+
## Objective
|
| 6 |
+
|
| 7 |
+
To build an email classification system for a support team that:
|
| 8 |
+
1. Masks PII (Full Name, Email, Phone, DOB, Aadhar, Card Numbers, CVV, Expiry) using Regex and SpaCy NER (without LLMs).
|
| 9 |
+
2. Classifies emails into categories (e.g., Billing Issues, Technical Support) using a trained machine learning model.
|
| 10 |
+
3. Exposes this functionality via a FastAPI endpoint.
|
| 11 |
|
| 12 |
## Project Structure
|
| 13 |
|
| 14 |
```
|
| 15 |
+
/workspaces/internship1/
|
| 16 |
+
├── saved_models/
|
| 17 |
+
│ └── email_classifier_pipeline.pkl # Saved classification model pipeline
|
| 18 |
+
├── api.py # FastAPI application logic and endpoints
|
| 19 |
+
├── app.py # Script to run the Uvicorn server
|
| 20 |
+
├── models.py # Model loading, prediction functions
|
| 21 |
+
├── utils.py # PII masking logic, text cleaning
|
| 22 |
+
├── train.py # Script to train the classification model
|
| 23 |
+
├── requirements.txt # Python package dependencies
|
| 24 |
+
├── combined_emails_with_natural_pii.csv # Dataset used for training (ensure this is present)
|
| 25 |
+
├── README.md # This file
|
| 26 |
+
└── .gitignore # (Optional but recommended: add *.pyc, __pycache__, .venv, saved_models/*, *.csv)
|
| 27 |
```
|
| 28 |
|
| 29 |
## Setup
|
| 30 |
|
| 31 |
+
1. **Clone the repository (if applicable):**
|
| 32 |
```bash
|
| 33 |
git clone <your-repo-url>
|
| 34 |
+
cd internship1
|
| 35 |
```
|
| 36 |
+
2. **Create and activate a virtual environment:**
|
|
|
|
| 37 |
```bash
|
| 38 |
+
python -m venv .venv
|
| 39 |
+
source .venv/bin/activate
|
| 40 |
+
# On Windows use: .venv\Scripts\activate
|
| 41 |
```
|
|
|
|
| 42 |
3. **Install dependencies:**
|
| 43 |
```bash
|
| 44 |
pip install -r requirements.txt
|
| 45 |
```
|
| 46 |
+
4. **Download SpaCy Model (if not done automatically by `utils.py`):**
|
|
|
|
| 47 |
```bash
|
| 48 |
python -m spacy download en_core_web_sm
|
|
|
|
|
|
|
| 49 |
```
|
| 50 |
+
5. **Train the Model (if `saved_models/email_classifier_pipeline.pkl` is not present):**
|
| 51 |
+
* Ensure the dataset (`combined_emails_with_natural_pii.csv`) is in the root directory.
|
| 52 |
+
* Run the training script:
|
| 53 |
+
```bash
|
| 54 |
+
python train.py
|
| 55 |
+
```
|
| 56 |
|
| 57 |
+
## Running the API
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 58 |
|
| 59 |
+
1. **Start the FastAPI server:**
|
| 60 |
+
```bash
|
| 61 |
+
python app.py
|
| 62 |
+
# OR directly using uvicorn:
|
| 63 |
+
# uvicorn api:app --host 0.0.0.0 --port 8000 --reload
|
| 64 |
+
```
|
| 65 |
+
The API will be available at `http://127.0.0.1:8000` (or the port specified).
|
|
|
|
|
|
|
| 66 |
|
| 67 |
+
## Using the API
|
| 68 |
|
| 69 |
+
Send a POST request to the `/classify/` endpoint with a JSON body containing the email text.
|
| 70 |
|
| 71 |
+
**Example using `curl`:**
|
| 72 |
|
| 73 |
+
```bash
|
| 74 |
+
curl -X POST "http://127.0.0.1:8000/classify/" \
|
| 75 |
+
-H "Content-Type: application/json" \
|
| 76 |
+
-d '{
|
| 77 |
+
"email_body": "Hello, my name is Jane Doe and my card number is 4111-1111-1111-1111. My email is jane.doe@example.com. Please help with billing."
|
| 78 |
+
}'
|
| 79 |
```
|
| 80 |
|
| 81 |
+
**Expected Response Structure:**
|
| 82 |
|
| 83 |
```json
|
| 84 |
{
|
| 85 |
+
"input_email_body": "Hello, my name is Jane Doe and my card number is 4111-1111-1111-1111. My email is jane.doe@example.com. Please help with billing.",
|
| 86 |
"list_of_masked_entities": [
|
| 87 |
{
|
| 88 |
+
"position": [18, 26],
|
| 89 |
"classification": "full_name",
|
| 90 |
+
"entity": "Jane Doe"
|
| 91 |
},
|
| 92 |
{
|
| 93 |
+
"position": [49, 68],
|
| 94 |
+
"classification": "credit_debit_no",
|
| 95 |
+
"entity": "4111-1111-1111-1111"
|
| 96 |
+
},
|
| 97 |
+
{
|
| 98 |
+
"position": [83, 103],
|
| 99 |
+
"classification": "email",
|
| 100 |
+
"entity": "jane.doe@example.com"
|
| 101 |
}
|
| 102 |
],
|
| 103 |
+
"masked_email": "Hello, my name is [full_name] and my card number is [credit_debit_no]. My email is [email]. Please help with billing.",
|
| 104 |
+
"category_of_the_email": "Billing Issues" // Example category
|
| 105 |
}
|
| 106 |
```
|
| 107 |
|
| 108 |
+
## Model Details
|
| 109 |
|
| 110 |
+
* **PII Masking:** SpaCy (`en_core_web_sm` for PERSON) and custom Regex patterns.
|
| 111 |
+
* **Classification Model:** TF-IDF Vectorizer + Multinomial Naive Bayes.
|
| 112 |
+
* **Training Accuracy:** Approx. 69% (as per last training run).
|
| 113 |
|
| 114 |
+
## Deployment
|
| 115 |
|
| 116 |
+
(Add link to your Hugging Face Space deployment here once completed)
|
| 117 |
+
|
| 118 |
+
```
|
| 119 |
+
HF Space: [Link]
|
| 120 |
+
```
|
__pycache__/api.cpython-312.pyc
CHANGED
|
Binary files a/__pycache__/api.cpython-312.pyc and b/__pycache__/api.cpython-312.pyc differ
|
|
|
__pycache__/models.cpython-312.pyc
CHANGED
|
Binary files a/__pycache__/models.cpython-312.pyc and b/__pycache__/models.cpython-312.pyc differ
|
|
|
api.py
CHANGED
|
@@ -1,20 +1,26 @@
|
|
| 1 |
from fastapi import FastAPI, HTTPException
|
| 2 |
from pydantic import BaseModel, Field
|
| 3 |
-
from typing import List, Dict, Tuple
|
| 4 |
-
import os
|
| 5 |
|
| 6 |
-
#
|
| 7 |
-
|
| 8 |
-
from
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 9 |
|
| 10 |
-
# ---
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
version="1.0.0"
|
| 15 |
-
)
|
| 16 |
|
| 17 |
-
# ---
|
| 18 |
class EmailInput(BaseModel):
|
| 19 |
email_body: str = Field(..., example="Hello, my name is Jane Doe and my email is jane.doe@example.com. I have a billing question.")
|
| 20 |
|
|
@@ -29,19 +35,23 @@ class ClassificationOutput(BaseModel):
|
|
| 29 |
masked_email: str
|
| 30 |
category_of_the_email: str
|
| 31 |
|
| 32 |
-
# --- Load
|
| 33 |
-
#
|
| 34 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
|
| 36 |
-
# --- API Endpoint ---
|
| 37 |
@app.post("/classify/", response_model=ClassificationOutput)
|
| 38 |
async def classify_email(email_input: EmailInput):
|
| 39 |
"""
|
| 40 |
Accepts an email body, masks PII, classifies the email,
|
| 41 |
and returns the results in the specified JSON format.
|
| 42 |
"""
|
| 43 |
-
if not
|
| 44 |
-
raise HTTPException(status_code=503, detail="Model
|
| 45 |
|
| 46 |
input_email = email_input.email_body
|
| 47 |
|
|
@@ -53,7 +63,7 @@ async def classify_email(email_input: EmailInput):
|
|
| 53 |
validated_entities = [MaskedEntity(**entity) for entity in entities]
|
| 54 |
|
| 55 |
# 2. Classify the masked email
|
| 56 |
-
predicted_class = predict_category(masked_email_body,
|
| 57 |
|
| 58 |
# 3. Construct the response
|
| 59 |
response = ClassificationOutput(
|
|
@@ -70,11 +80,6 @@ async def classify_email(email_input: EmailInput):
|
|
| 70 |
# Optionally include more detail in the error response during development
|
| 71 |
raise HTTPException(status_code=500, detail=f"Internal Server Error: {str(e)}")
|
| 72 |
|
| 73 |
-
# --- Root Endpoint (Optional - for basic check) ---
|
| 74 |
-
@app.get("/")
|
| 75 |
-
async def root():
|
| 76 |
-
return {"message": "Email Classification API is running. Use the /classify/ endpoint."}
|
| 77 |
-
|
| 78 |
# --- Running the API (for local development) ---
|
| 79 |
# You'll typically run this using uvicorn from the terminal:
|
| 80 |
# uvicorn api:app --reload
|
|
|
|
| 1 |
from fastapi import FastAPI, HTTPException
|
| 2 |
from pydantic import BaseModel, Field
|
| 3 |
+
from typing import List, Dict, Tuple, Any, Optional
|
| 4 |
+
import os # To potentially load environment variables if needed
|
| 5 |
|
| 6 |
+
# Make sure these imports work and don't cause errors themselves
|
| 7 |
+
try:
|
| 8 |
+
from utils import mask_pii
|
| 9 |
+
from models import load_model_pipeline, predict_category, Pipeline # Ensure Pipeline is imported if type hint used
|
| 10 |
+
except ImportError as e:
|
| 11 |
+
print(f"Error importing from utils or models in api.py: {e}")
|
| 12 |
+
# Optionally raise the error to make it obvious during startup
|
| 13 |
+
# raise e
|
| 14 |
+
except Exception as e:
|
| 15 |
+
print(f"Unexpected error during imports in api.py: {e}")
|
| 16 |
+
# raise e
|
| 17 |
|
| 18 |
+
# --- FastAPI App ---
|
| 19 |
+
# >>>>> THIS LINE IS CRUCIAL <<<<<
|
| 20 |
+
app = FastAPI()
|
| 21 |
+
# >>>>> MUST BE NAMED 'app' <<<<<
|
|
|
|
|
|
|
| 22 |
|
| 23 |
+
# --- Pydantic Models ---
|
| 24 |
class EmailInput(BaseModel):
|
| 25 |
email_body: str = Field(..., example="Hello, my name is Jane Doe and my email is jane.doe@example.com. I have a billing question.")
|
| 26 |
|
|
|
|
| 35 |
masked_email: str
|
| 36 |
category_of_the_email: str
|
| 37 |
|
| 38 |
+
# --- Load Model at Startup ---
|
| 39 |
+
# Ensure load_model_pipeline() exists and works
|
| 40 |
+
model_pipeline: Optional[Pipeline] = load_model_pipeline()
|
| 41 |
+
|
| 42 |
+
# --- API Endpoints ---
|
| 43 |
+
@app.get("/")
|
| 44 |
+
async def read_root():
|
| 45 |
+
return {"message": "Email Classification API is running. Use the /classify/ endpoint."}
|
| 46 |
|
|
|
|
| 47 |
@app.post("/classify/", response_model=ClassificationOutput)
|
| 48 |
async def classify_email(email_input: EmailInput):
|
| 49 |
"""
|
| 50 |
Accepts an email body, masks PII, classifies the email,
|
| 51 |
and returns the results in the specified JSON format.
|
| 52 |
"""
|
| 53 |
+
if not model_pipeline:
|
| 54 |
+
raise HTTPException(status_code=503, detail="Model pipeline not available. Please train/load it first.")
|
| 55 |
|
| 56 |
input_email = email_input.email_body
|
| 57 |
|
|
|
|
| 63 |
validated_entities = [MaskedEntity(**entity) for entity in entities]
|
| 64 |
|
| 65 |
# 2. Classify the masked email
|
| 66 |
+
predicted_class = predict_category(masked_email_body, model_pipeline)
|
| 67 |
|
| 68 |
# 3. Construct the response
|
| 69 |
response = ClassificationOutput(
|
|
|
|
| 80 |
# Optionally include more detail in the error response during development
|
| 81 |
raise HTTPException(status_code=500, detail=f"Internal Server Error: {str(e)}")
|
| 82 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 83 |
# --- Running the API (for local development) ---
|
| 84 |
# You'll typically run this using uvicorn from the terminal:
|
| 85 |
# uvicorn api:app --reload
|
app.py
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# filepath: /workspaces/internship1/app.py
|
| 2 |
+
import uvicorn
|
| 3 |
+
import os
|
| 4 |
+
|
| 5 |
+
# Import the FastAPI app instance from api.py
|
| 6 |
+
# Ensure the FastAPI instance in api.py is named 'app'
|
| 7 |
+
try:
|
| 8 |
+
from api import app
|
| 9 |
+
except ImportError:
|
| 10 |
+
print("Error: Could not import 'app' from api.py.")
|
| 11 |
+
print("Make sure api.py exists and contains a FastAPI instance named 'app'.")
|
| 12 |
+
app = None
|
| 13 |
+
except Exception as e:
|
| 14 |
+
print(f"An unexpected error occurred during import: {e}")
|
| 15 |
+
app = None
|
| 16 |
+
|
| 17 |
+
|
| 18 |
+
if __name__ == "__main__":
|
| 19 |
+
if app:
|
| 20 |
+
# Get port from environment variable PORT, default to 8000
|
| 21 |
+
# Hugging Face Spaces and other platforms often set the PORT variable
|
| 22 |
+
port = int(os.environ.get("PORT", 8000))
|
| 23 |
+
# Use host="0.0.0.0" to make it accessible externally (in Codespaces/Docker/HF)
|
| 24 |
+
print(f"Starting Uvicorn server on host 0.0.0.0, port {port}")
|
| 25 |
+
uvicorn.run(app, host="0.0.0.0", port=port)
|
| 26 |
+
else:
|
| 27 |
+
print("Could not start server because the FastAPI app instance was not loaded.")
|
generate_output.py
ADDED
|
@@ -0,0 +1,100 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import pandas as pd
|
| 2 |
+
import requests
|
| 3 |
+
import json
|
| 4 |
+
from pathlib import Path
|
| 5 |
+
import time
|
| 6 |
+
|
| 7 |
+
# --- Configuration ---
|
| 8 |
+
DATASET_PATH = Path("combined_emails_with_natural_pii.csv") # Path to your input CSV
|
| 9 |
+
OUTPUT_PATH = Path("api_output_results.jsonl") # Where to save the results (JSON Lines format)
|
| 10 |
+
API_ENDPOINT = "http://127.0.0.1:8000/classify/" # Your running API endpoint
|
| 11 |
+
|
| 12 |
+
# Adjust column names if different
|
| 13 |
+
email_body_column = 'email'
|
| 14 |
+
# category_column = 'type' # Original category column (optional, for reference)
|
| 15 |
+
|
| 16 |
+
# --- Main Function ---
|
| 17 |
+
def process_emails_via_api(data_path: Path, output_path: Path, api_url: str):
|
| 18 |
+
"""
|
| 19 |
+
Reads emails from a CSV, sends them to the classification API,
|
| 20 |
+
and saves the JSON responses to a file.
|
| 21 |
+
"""
|
| 22 |
+
if not data_path.exists():
|
| 23 |
+
print(f"Error: Input dataset not found at {data_path}")
|
| 24 |
+
return
|
| 25 |
+
|
| 26 |
+
print(f"Loading dataset from {data_path}...")
|
| 27 |
+
try:
|
| 28 |
+
# Skip bad lines just in case, consistent with training
|
| 29 |
+
df = pd.read_csv(data_path, on_bad_lines='skip')
|
| 30 |
+
except Exception as e:
|
| 31 |
+
print(f"Error loading CSV: {e}")
|
| 32 |
+
return
|
| 33 |
+
|
| 34 |
+
if email_body_column not in df.columns:
|
| 35 |
+
print(f"Error: Email body column '{email_body_column}' not found.")
|
| 36 |
+
return
|
| 37 |
+
|
| 38 |
+
# Handle potential missing email bodies
|
| 39 |
+
df.dropna(subset=[email_body_column], inplace=True)
|
| 40 |
+
if df.empty:
|
| 41 |
+
print("Error: No valid email bodies found after handling missing values.")
|
| 42 |
+
return
|
| 43 |
+
|
| 44 |
+
results = []
|
| 45 |
+
total_emails = len(df)
|
| 46 |
+
print(f"Processing {total_emails} emails via API: {api_url}")
|
| 47 |
+
|
| 48 |
+
# Open output file in write mode (clears existing content)
|
| 49 |
+
with open(output_path, 'w') as f_out:
|
| 50 |
+
for index, row in df.iterrows():
|
| 51 |
+
email_text = str(row[email_body_column]) # Ensure it's a string
|
| 52 |
+
payload = {"email_body": email_text}
|
| 53 |
+
|
| 54 |
+
try:
|
| 55 |
+
response = requests.post(api_url, json=payload, timeout=30) # Added timeout
|
| 56 |
+
response.raise_for_status() # Raise an exception for bad status codes (4xx or 5xx)
|
| 57 |
+
|
| 58 |
+
api_result = response.json()
|
| 59 |
+
|
| 60 |
+
# Write result as a JSON line
|
| 61 |
+
f_out.write(json.dumps(api_result) + '\n')
|
| 62 |
+
|
| 63 |
+
if (index + 1) % 50 == 0: # Print progress every 50 emails
|
| 64 |
+
print(f"Processed {index + 1}/{total_emails} emails...")
|
| 65 |
+
|
| 66 |
+
except requests.exceptions.RequestException as e:
|
| 67 |
+
print(f"\nError processing email index {index}: {e}")
|
| 68 |
+
# Optionally write error info to the file or a separate log
|
| 69 |
+
error_info = {
|
| 70 |
+
"error": str(e),
|
| 71 |
+
"input_email_body": email_text,
|
| 72 |
+
"index": index
|
| 73 |
+
}
|
| 74 |
+
f_out.write(json.dumps(error_info) + '\n')
|
| 75 |
+
# Optional: add a small delay if API is overloaded
|
| 76 |
+
# time.sleep(0.5)
|
| 77 |
+
except json.JSONDecodeError as e:
|
| 78 |
+
print(f"\nError decoding JSON response for email index {index}: {e}")
|
| 79 |
+
print(f"Response status code: {response.status_code}")
|
| 80 |
+
print(f"Response text: {response.text[:500]}...") # Print beginning of text
|
| 81 |
+
error_info = {
|
| 82 |
+
"error": f"JSONDecodeError: {e}",
|
| 83 |
+
"response_text": response.text,
|
| 84 |
+
"input_email_body": email_text,
|
| 85 |
+
"index": index
|
| 86 |
+
}
|
| 87 |
+
f_out.write(json.dumps(error_info) + '\n')
|
| 88 |
+
|
| 89 |
+
|
| 90 |
+
print(f"\nProcessing complete. Results saved to {output_path}")
|
| 91 |
+
|
| 92 |
+
|
| 93 |
+
# --- Script Execution ---
|
| 94 |
+
if __name__ == "__main__":
|
| 95 |
+
# Make sure the API is running before executing this script!
|
| 96 |
+
print("--- Starting API Output Generation ---")
|
| 97 |
+
print("Ensure the FastAPI server (python app.py or uvicorn) is running in another terminal.")
|
| 98 |
+
input("Press Enter to continue once the API is running...")
|
| 99 |
+
process_emails_via_api(DATASET_PATH, OUTPUT_PATH, API_ENDPOINT)
|
| 100 |
+
print("--- Finished API Output Generation ---")
|
models.py
CHANGED
|
@@ -10,7 +10,6 @@ import re
|
|
| 10 |
from fastapi import FastAPI, HTTPException
|
| 11 |
from pydantic import BaseModel
|
| 12 |
from utils import clean_text_for_classification, mask_pii
|
| 13 |
-
from models import MODEL_PATH, load_model_pipeline, predict_category
|
| 14 |
|
| 15 |
# --- Constants ---
|
| 16 |
MODEL_DIR = Path("saved_models")
|
|
@@ -37,7 +36,7 @@ class ClassificationOutput(BaseModel):
|
|
| 37 |
|
| 38 |
# --- Load Model at Startup ---
|
| 39 |
# Load the model pipeline once when the application starts
|
| 40 |
-
model_pipeline: Optional[Pipeline] =
|
| 41 |
|
| 42 |
# --- Model Loading ---
|
| 43 |
def load_model_pipeline() -> Optional[Pipeline]:
|
|
|
|
| 10 |
from fastapi import FastAPI, HTTPException
|
| 11 |
from pydantic import BaseModel
|
| 12 |
from utils import clean_text_for_classification, mask_pii
|
|
|
|
| 13 |
|
| 14 |
# --- Constants ---
|
| 15 |
MODEL_DIR = Path("saved_models")
|
|
|
|
| 36 |
|
| 37 |
# --- Load Model at Startup ---
|
| 38 |
# Load the model pipeline once when the application starts
|
| 39 |
+
model_pipeline: Optional[Pipeline] = None
|
| 40 |
|
| 41 |
# --- Model Loading ---
|
| 42 |
def load_model_pipeline() -> Optional[Pipeline]:
|
requirements.txt
CHANGED
|
@@ -1,19 +1,19 @@
|
|
| 1 |
-
# Core
|
| 2 |
-
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
|
|
|
|
|
|
| 6 |
|
| 7 |
-
#
|
| 8 |
-
spacy
|
| 9 |
-
|
|
|
|
| 10 |
|
| 11 |
-
#
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
pydantic>=1.8.0,<2.0.0 # For data validation in FastAPI
|
| 15 |
|
| 16 |
-
#
|
| 17 |
-
#
|
| 18 |
-
# joblib or pickle for saving models if not using framework methods
|
| 19 |
-
joblib
|
|
|
|
| 1 |
+
# --- Core ---
|
| 2 |
+
fastapi==0.111.0
|
| 3 |
+
uvicorn[standard]==0.29.0
|
| 4 |
+
pandas==2.2.2
|
| 5 |
+
numpy==1.26.4
|
| 6 |
+
scikit-learn==1.4.2
|
| 7 |
+
joblib==1.4.2
|
| 8 |
|
| 9 |
+
# --- PII Masking ---
|
| 10 |
+
spacy==3.7.4
|
| 11 |
+
# The following model is downloaded by spacy CLI, but good to list
|
| 12 |
+
# en_core_web_sm @ https://github.com/explosion/spacy-models/releases/download/en_core_web_sm-3.7.1/en_core_web_sm-3.7.1-py3-none-any.whl
|
| 13 |
|
| 14 |
+
# --- Other potential dependencies (add if you used them) ---
|
| 15 |
+
# nltk==3.8.1 # If you used NLTK for preprocessing
|
| 16 |
+
# python-dotenv==1.0.1 # If loading environment variables
|
|
|
|
| 17 |
|
| 18 |
+
# --- Development/Testing (Optional) ---
|
| 19 |
+
# requests==2.31.0 # For testing the API or running generate_output.py
|
|
|
|
|
|