metascraper
PackageNode.js library that extracts a page's title, image and description from HTML
- Price
- Free, open source
- Access
- None, runs locally
About
metascraper is an MIT-licensed Node.js library that extracts unified metadata (title, description, image, author, date, logo, publisher) from a page's HTML with Open Graph, Microdata, RDFa, Twitter Card and JSON-LD rules. You fetch the HTML yourself. It powers Microlink; Node 22+.
What you can do with it
- Build your own URL unfurler from a page's Open Graph and JSON-LD tags
- Extract title, author, date and lead image from an article URL in Node.js
- Add custom rules that pull extra metadata fields from fetched HTML
Get started
- Run npm i metascraper and rule packages like metascraper-title
- Fetch the page HTML, then pass html and url to metascraper
Example
import createMetascraper from "metascraper";
import description from "metascraper-description";
import image from "metascraper-image";
import title from "metascraper-title";
const metascraper = createMetascraper([title(), description(), image()]);
const url = "https://example.com/article";
const html = await (await fetch(url)).text();
console.log(await metascraper({ html, url }));Details
- Hosting
- Runs locally
- Available in
- Worldwide
- Official SDKs
- JavaScript/TypeScript
- MCP server
- None