metascraper

Package

Node.js library that extracts a page's title, image and description from HTML

Price
Free, open source
Access
None, runs locally

About

metascraper is an MIT-licensed Node.js library that extracts unified metadata (title, description, image, author, date, logo, publisher) from a page's HTML with Open Graph, Microdata, RDFa, Twitter Card and JSON-LD rules. You fetch the HTML yourself. It powers Microlink; Node 22+.

What you can do with it

  • Build your own URL unfurler from a page's Open Graph and JSON-LD tags
  • Extract title, author, date and lead image from an article URL in Node.js
  • Add custom rules that pull extra metadata fields from fetched HTML

Get started

  1. Run npm i metascraper and rule packages like metascraper-title
  2. Fetch the page HTML, then pass html and url to metascraper

Example

import createMetascraper from "metascraper";
import description from "metascraper-description";
import image from "metascraper-image";
import title from "metascraper-title";

const metascraper = createMetascraper([title(), description(), image()]);
const url = "https://example.com/article";
const html = await (await fetch(url)).text();
console.log(await metascraper({ html, url }));

Details

Hosting
Runs locally
Available in
Worldwide
Official SDKs
JavaScript/TypeScript
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .