---
source_url: "https://agenty.com/tools/content?utm_source=openai"
title: "Article Content Extraction API - Clean HTML & Text - Agenty"
mirrored_at: 2026-08-07T01:08:48.727Z
host: agenty.com
cited_in_42a: true
mirror_canonical: "https://index.42a.ai/agenty.com/tools/content__q__utm_source_openai"
---

> **Original source:** https://agenty.com/tools/content?utm_source=openai

1.  [Home](https://agenty.com/)
2.  [Tools](https://agenty.com/tools)
3.  Article Content Extraction API

Extract clean article body, title, author, and publish date from any blog or news page — without the ads and clutter.

The Agenty Content API extracts the main article body from any blog or news URL, automatically removing navigation, ads, sidebars, and footers. Get the title, author, publish date, hero image, and a clean HTML or plain-text body — ready for aggregators, newsletters, or [LLM pipelines](https://agenty.com/docs/mcp-server-web-scraping-automation-with-ai/536).

## Features

## Use cases

-   News and blog aggregation feeds
-   Newsletter and content curation pipelines
-   Competitive content and SEO analysis
-   Building clean text corpora for LLM training
-   Reader-mode features in apps and extensions

## API examples

## How Agenty compares

Feature

Agenty

Readability

Mercury

Postlight

Automatic content detection

Yes

Yes

Yes

Yes

Author & date extraction

Yes

Limited

Yes

Yes

Multi-language (50+)

Yes

Limited

Yes

Yes

Image & video extraction

Yes

No

Yes

Yes

Hosted API + free tier

Yes

Self-host

Self-host

Yes

## Frequently asked questions

### What is the Article Content Extraction API?

The Agenty Content API automatically identifies and extracts the main article on any web page. It returns clean structured data: title, author, publish date, article body, and embedded media URLs.

### Can I get plain text instead of HTML?

Yes. Set `outputFormat: "text"` to receive plain text with paragraphs preserved. The default is `"html"` which returns clean semantic HTML.

### Does it work with paywalled content?

Yes. Pass cookies or auth headers via the `headers` parameter. We also support session-based authentication for platforms like Medium and Substack.

### Is there a free tier?

Yes. All accounts include a free tier. Visit our [pricing page](https://agenty.com/pricing) for details.

## Explore more

-   [Pricing](https://agenty.com/pricing)
-   [Documentation](https://agenty.com/docs)
-   [Website Screenshot API](https://agenty.com/tools/screenshot)
-   [Web Data Extraction API](https://agenty.com/tools/extract)
-   [HTML to Markdown API](https://agenty.com/tools/markdown)
-   [HTML to PDF API](https://agenty.com/tools/pdf)
-   [Web Scraping API](https://agenty.com/tools/scrape)
-   [URL Redirect Capture API](https://agenty.com/tools/redirects)
-   [Website Link Extractor API](https://agenty.com/tools/links)