Page Pulse is a Spring Boot REST API that audits a web page by accepting a URL and returning a JSON report containing useful information about the page.
The API performs the following checks:
- HTTP Status Code
- Response Time
- Page Title
- Meta Description
- Number of H1 Tags
- Number of Images Missing Alt Text
- Approximate Word Count
The application also handles common failure scenarios such as invalid URLs, request timeouts, non-HTML resources, and network errors without crashing.
- Java 21
- Spring Boot
- Maven
- Jsoup
- JUnit 5
src
├── controller
├── dto
├── exception
├── service
│ ├── AuditService
│ ├── AuditServiceImpl
│ └── HtmlParser
- Java 21 or later
- Maven 3.9+
git clone <repository-url>mvn spring-boot:runThe application will start on:
http://localhost:8080
Endpoint
POST /api/audit
{
"url": "https://google.com"
}{
"httpStatus": 200,
"responseTime": 243,
"pageTitle": "Google",
"metaDescription": "Search the world's information...",
"h1Count": 1,
"imagesMissingAlt": 0,
"approximateWordCount": 342
}{
"message": "Invalid URL"
}{
"message": "Request timed out."
}{
"message": "URL does not point to an HTML page."
}{
"message": "Unable to fetch the webpage."
}Unit tests have been written for the HTML parsing logic.
The tests cover:
- Happy path parsing
- Missing meta description
- Images without alt text
Run the tests using:
mvn testor directly from IntelliJ using the JUnit test runner.
The application separates webpage retrieval from HTML parsing by introducing a dedicated HtmlParser component.
Reason
This follows the Single Responsibility Principle and makes the parsing logic independently testable without requiring network requests.
Custom exceptions together with a global exception handler are used to return consistent JSON error responses.
Reason
This keeps the service layer focused on business logic while ensuring clients receive meaningful and predictable error messages.
Jsoup is used for both downloading HTML pages and parsing their contents.
Reason
Jsoup provides a simple API for HTTP requests and DOM traversal, allowing the application to extract information such as titles, headings, images, and meta tags with minimal code.
If given additional time, the following enhancements could be implemented:
- Detect broken links on the page.
- Analyze heading hierarchy (H1, H2, H3).
- Improve accessibility checks beyond missing alt attributes.
- Support asynchronous page fetching for improved performance.
- Cache audit results to reduce repeated network requests.
- Add authentication and rate limiting for production deployments.
Developed as part of the Page Pulse backend assessment using Spring Boot and Jsoup.