Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Link Parser Backend

Overview

Page Pulse is a Spring Boot REST API that audits a web page by accepting a URL and returning a JSON report containing useful information about the page.

The API performs the following checks:

  • HTTP Status Code
  • Response Time
  • Page Title
  • Meta Description
  • Number of H1 Tags
  • Number of Images Missing Alt Text
  • Approximate Word Count

The application also handles common failure scenarios such as invalid URLs, request timeouts, non-HTML resources, and network errors without crashing.


Tech Stack

  • Java 21
  • Spring Boot
  • Maven
  • Jsoup
  • JUnit 5

Project Structure

src
├── controller
├── dto
├── exception
├── service
│   ├── AuditService
│   ├── AuditServiceImpl
│   └── HtmlParser

Setup

Prerequisites

  • Java 21 or later
  • Maven 3.9+

Clone the repository

git clone <repository-url>

Run the application

mvn spring-boot:run

The application will start on:

http://localhost:8080

API Contract

Audit a Web Page

Endpoint

POST /api/audit

Request Body

{
    "url": "https://google.com"
}

Successful Response

{
    "httpStatus": 200,
    "responseTime": 243,
    "pageTitle": "Google",
    "metaDescription": "Search the world's information...",
    "h1Count": 1,
    "imagesMissingAlt": 0,
    "approximateWordCount": 342
}

Error Responses

Invalid URL

{
    "message": "Invalid URL"
}

Request Timeout

{
    "message": "Request timed out."
}

Non-HTML Response

{
    "message": "URL does not point to an HTML page."
}

Unable to Fetch Web Page

{
    "message": "Unable to fetch the webpage."
}

Testing

Unit tests have been written for the HTML parsing logic.

The tests cover:

  • Happy path parsing
  • Missing meta description
  • Images without alt text

Run the tests using:

mvn test

or directly from IntelliJ using the JUnit test runner.


Design Decisions

1. Separated Fetching and Parsing Logic

The application separates webpage retrieval from HTML parsing by introducing a dedicated HtmlParser component.

Reason

This follows the Single Responsibility Principle and makes the parsing logic independently testable without requiring network requests.


2. Global Exception Handling

Custom exceptions together with a global exception handler are used to return consistent JSON error responses.

Reason

This keeps the service layer focused on business logic while ensuring clients receive meaningful and predictable error messages.


3. Jsoup for HTML Processing

Jsoup is used for both downloading HTML pages and parsing their contents.

Reason

Jsoup provides a simple API for HTTP requests and DOM traversal, allowing the application to extract information such as titles, headings, images, and meta tags with minimal code.


Future Improvements

If given additional time, the following enhancements could be implemented:

  • Detect broken links on the page.
  • Analyze heading hierarchy (H1, H2, H3).
  • Improve accessibility checks beyond missing alt attributes.
  • Support asynchronous page fetching for improved performance.
  • Cache audit results to reduce repeated network requests.
  • Add authentication and rate limiting for production deployments.

Author

Developed as part of the Page Pulse backend assessment using Spring Boot and Jsoup.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages