Skip to content

About

Tamamen yerel çalışan AI metin insanlaştırıcı — Türkçe ve İngilizce, Ollama ile çevrimdışı LLM çıkarımı

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Text Humanizer

License: MIT Python Tests

A fully local AI text humanizer that converts AI-generated text into human-sounding writing. Supports English and Turkish. Works offline using Ollama for LLM inference; all rule-based modes require no network connection.


Features

  • Rule-based pipeline (Levels 1-2): fast, deterministic, no GPU required
  • LLM pipeline (Levels 3-4): uses a local Ollama model for deeper rewriting
  • Language support: English and Turkish (including Turkish formal suffix reduction and vowel harmony)
  • Benchmark metrics: burstiness, lexical diversity, entropy, AI phrase density, readability grade, passive voice density, formal marker density, humanness score
  • Streamlit GUI and Typer CLI
  • File input: plain text (.txt), Word (.docx), PDF (.pdf)

Screenshots

Main interface — sidebar settings, original text input, humanized output, and benchmark score summary:

Main interface

Benchmark panel — detailed metrics table (before/after/delta/target) and radar chart:

Benchmark panel

Sentence analysis — top AI-scored sentences with alternative suggestions:

Sentence analysis


Requirements

  • Python 3.10 or higher
  • Ollama installed and running (required for Levels 3-4 only)
  • Windows, macOS, or Linux

Installation

1. Clone the repository

git clone https://github.com/Yukseltt/text-humanizer.git
cd text-humanizer

2. Create and activate a virtual environment

Windows (PowerShell):

python -m venv venv
.\venv\Scripts\Activate.ps1

macOS / Linux:

python -m venv venv
source venv/bin/activate

3. Install dependencies

pip install -r requirements.txt

4. Download NLTK data

python -c "import nltk; nltk.download('punkt_tab')"

5. Configure environment

Copy the example file and edit it:

cp .env.example .env

Edit .env:

OLLAMA_HOST=http://localhost:11434
HUMANIZER_MODEL=llama3.2:3b
CHUNK_SIZE=600

Model options by VRAM:

  • llama3.2:3b — 2 GB, works on 4-6 GB VRAM (default)
  • llama3.1:8b — 4.7 GB, works on 8 GB+ VRAM (better quality)
  • mistral:7b — 4.1 GB, alternative for 8 GB+ VRAM

6. Pull the Ollama model (required for Levels 3-4)

ollama pull llama3.2:3b

Running

GUI (Streamlit) — Windows

Two batch files are provided for one-click startup on Windows.

Step 1 — Start the Ollama server (required for Levels 3-4):

Double-click start_ollama.bat (or run the PowerShell script directly):

start_ollama.bat

What it does:

  • Checks that Ollama is installed (%LOCALAPPDATA%\Programs\Ollama\ollama.exe)
  • Detects whether port 11434 is already in use and offers to kill the existing process
  • Starts ollama serve and keeps the window open
  • Shuts down Ollama cleanly when the window is closed

Keep this window open while using the GUI.

Step 2 — Launch the GUI:

Double-click run_gui.bat (or the PowerShell equivalent):

run_gui.bat

What it does:

  • Activates the venv environment
  • Runs streamlit run app.py
  • Prints the local URL in the terminal

The GUI opens at http://localhost:8501.

PowerShell equivalents:

PowerShell -ExecutionPolicy Bypass -File start_ollama.ps1
PowerShell -ExecutionPolicy Bypass -File run_gui.ps1

Manual start (any OS):

ollama serve          # terminal 1 — keep open
streamlit run app.py  # terminal 2

CLI

Humanize text directly:

python -m src.cli "Furthermore, it is important to note that this is comprehensive."

Humanize from file:

python -m src.cli --input input.txt --output result.txt

Rules only (no Ollama required):

python -m src.cli --no-ollama --input input.txt

With benchmark report:

python -m src.cli --benchmark --input input.txt

Standalone benchmark between two files:

python -m src.cli benchmark input.txt --humanized result.txt

Tests

pytest tests/ -v

All 118 tests should pass in under 5 seconds.


Project Structure

text-humanizer/
├── app.py                      # Streamlit GUI
├── config.py                   # Pydantic settings (reads .env)
├── requirements.txt
├── .env.example
├── start_ollama.bat            # Windows: start Ollama server (double-click)
├── start_ollama.ps1            # PowerShell: start Ollama server
├── run_gui.bat                 # Windows: launch Streamlit GUI (double-click)
├── run_gui.ps1                 # PowerShell: launch Streamlit GUI
│
├── src/
│   ├── cli.py                  # Typer CLI entry point
│   ├── humanizer/
│   │   ├── core.py             # Main humanization pipeline
│   │   ├── techniques.py       # Rule-based transforms
│   │   ├── preprocessor.py     # Text splitting, normalization, header detection
│   │   ├── analyzer.py         # Sentence AI scoring, entropy, references
│   │   ├── prompt.py           # Ollama prompt templates
│   │   └── file_handler.py     # .txt / .docx / .pdf reading
│   └── benchmark/
│       ├── metrics.py          # All metric calculations
│       ├── runner.py           # BenchmarkResult class
│       └── report.py          # Rich terminal report
│
├── data/
│   └── collocations.json       # EN and TR collocation replacement table
│
└── tests/
    ├── conftest.py
    ├── test_metrics.py
    ├── test_techniques.py
    ├── test_analyzer.py
    ├── test_preprocessor.py
    └── test_benchmark_runner.py

Humanization Levels

Level Mode Description
1 Rules AI phrase replacement, contractions (EN), Turkish formal suffix reduction
2 Rules Level 1 + collocation replacement (longest-match, EN and TR)
3 LLM Level 1-2 rules, then Ollama LLM rewrite
4 LLM Level 3 + multi-pass (rules -> LLM -> rules)

Levels 1-2 are offline and deterministic. Levels 3-4 require Ollama.


Benchmark Metrics

Metric What it measures
Humanness Score Composite score (0-100), higher is more human-like
AI Phrase Density Percentage of AI filler phrases (lower is better)
Burstiness Variance in sentence lengths (higher = more human rhythm)
Lexical Diversity Type-token ratio (higher = more varied vocabulary)
Entropy Shannon entropy of word distribution (higher = more diverse)
Readability Grade Flesch-Kincaid grade level (target ~10 for accessible writing)
Passive Voice Density Percentage of passive sentences (lower is better)
Formal Marker Density Percentage of formal marker words (lower is better)

Configuration Reference

All settings are read from .env via config.py.

Variable Default Description
OLLAMA_HOST http://localhost:11434 Ollama server address
HUMANIZER_MODEL llama3.2:3b Model to use for Levels 3-4
CHUNK_SIZE 600 Max words per LLM chunk for long texts

Troubleshooting

Ollama connection error: Start Ollama first with start_ollama.bat (Windows) or ollama serve (terminal), then launch the GUI. Check that OLLAMA_HOST in .env matches (http://localhost:11434).

Model not found: Run ollama pull <model-name> before using Level 3-4.

NLTK punkt not found: Run python -c "import nltk; nltk.download('punkt_tab')".

Streamlit port conflict: Run streamlit run app.py --server.port 8502.

About

Tamamen yerel çalışan AI metin insanlaştırıcı — Türkçe ve İngilizce, Ollama ile çevrimdışı LLM çıkarımı

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages