metascraper

Paket

Node.js library that extracts a page's title, image and description from HTML

Fiyat
Free, open source
Erişim
None, runs locally

Hakkında

metascraper is an MIT-licensed Node.js library that extracts unified metadata (title, description, image, author, date, logo, publisher) from a page's HTML with Open Graph, Microdata, RDFa, Twitter Card and JSON-LD rules. You fetch the HTML yourself. It powers Microlink; Node 22+.

Neler yapabilirsin

  • Build your own URL unfurler from a page's Open Graph and JSON-LD tags
  • Extract title, author, date and lead image from an article URL in Node.js
  • Add custom rules that pull extra metadata fields from fetched HTML

Başlarken

  1. Run npm i metascraper and rule packages like metascraper-title
  2. Fetch the page HTML, then pass html and url to metascraper

Örnek kod

import createMetascraper from "metascraper";
import description from "metascraper-description";
import image from "metascraper-image";
import title from "metascraper-title";

const metascraper = createMetascraper([title(), description(), image()]);
const url = "https://example.com/article";
const html = await (await fetch(url)).text();
console.log(await metascraper({ html, url }));

Ayrıntılar

Barındırma
Kendi bilgisayarında
Kullanılabildiği yerler
Tüm dünya
Resmi SDK'lar
JavaScript/TypeScript
MCP sunucusu
Yok

Görevler

Alternatifler

Aynı görevler için başka araçlar.

Son kontrol: .