npmpackage.info

Gathering detailed insights and metrics for @crawlee/cheerio

Other packages similar to @crawlee/cheerio

@vladfrangu-dev/crawlee-cheerio

3.3.0

The scalable web crawling and scraping library for JavaScript/Node.js. Enables development of data extraction and web automation jobs (not only) with headless Chrome and Puppeteer.

Gathering detailed insights and metrics for @crawlee/cheerio

@crawlee/cheerio - 3.13.10 | npmpackage.info

@crawlee/cheerio

Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.

3.13.10

18,480

Apache-2.0

TypeScript

210.71 kB

Installations

npm install @crawlee/cheerio

Developer Guide

BETA

Typescript

Yes

Module System

CommonJS, ESM

Min. Node Version

>=16.0.0

Node Version

20.19.3

NPM Version

lerna/8.2.3/node@v20.19.3+x64 (linux)

Pull Requests

Open

18

Total

1,729

Closed

378

Merged

1,333

Issues

Open

129

Total

964

Closed

835

Releases

126

v3.13.10

Updated on Jul 09, 2025

v3.13.9

Updated on Jun 27, 2025

v3.13.8

Updated on Jun 16, 2025

v3.13.7

Updated on Jun 06, 2025

v3.13.6

Updated on Jun 05, 2025

v3.13.5

Updated on May 20, 2025

View All 126 releases

Languages

TypeScript

MDX

JavaScript

CSS

Dockerfile

Python

HTML

TypeScript (49.05%)

MDX (44.29%)

JavaScript (5.14%)

CSS (0.84%)

Dockerfile (0.67%)

Python (0.02%)

Developer

apify

Download Statistics

Total Downloads

Last Day

Last Week

Last Month

Last Year

GitHub Statistics

Apache-2.0 License

18,480 Stars

4,990 Commits

884 Forks

113 Watchers

25 Branches

110 Contributors

Updated on Jul 15, 2025

Maintainers

View All 110 Contributors

Package Meta Information

Latest Version

3.13.10

Package Id

@crawlee/cheerio@3.13.10

Unpacked Size

210.71 kB

Size

62.25 kB

File Count

NPM Version

lerna/8.2.3/node@v20.19.3+x64 (linux)

Node Version

20.19.3

Published on

Jul 09, 2025

Total Downloads

Cumulative downloads

Total Downloads

NaN

Last Day

NaN

Compared to previous day

Last Week

NaN

Compared to previous week

Last Month

NaN

Compared to previous month

Last Year

NaN

Compared to previous year

Weekly Downloads

Monthly Downloads

Yearly Downloads

Dependencies

@crawlee/http @crawlee/types @crawlee/utils cheerio htmlparser2 tslib

A web scraping and browser automation library

Crawlee covers your crawling and scraping end-to-end and helps you build reliable scrapers. Fast.

Your crawlers will appear human-like and fly under the radar of modern bot protections even with the default configuration. Crawlee gives you the tools to crawl the web for links, scrape data, and store it to disk or cloud while staying configurable to suit your project's needs.

Crawlee is available as the crawlee NPM package.

👉 View full documentation, guides and examples on the Crawlee project website 👈

Crawlee for Python is open for early adopters. 🐍 👉 Checkout the source code 👈.

Installation

We recommend visiting the Introduction tutorial in Crawlee documentation for more information.

Crawlee requires Node.js 16 or higher.

With Crawlee CLI

The fastest way to try Crawlee out is to use the Crawlee CLI and choose the Getting started example. The CLI will install all the necessary dependencies and add boilerplate code for you to play with.

1npx crawlee create my-crawler

1cd my-crawler
2npm start

Manual installation

If you prefer adding Crawlee into your own project, try the example below. Because it uses PlaywrightCrawler we also need to install Playwright. It's not bundled with Crawlee to reduce install size.

1npm install crawlee playwright

1import { PlaywrightCrawler, Dataset } from 'crawlee';
2
3// PlaywrightCrawler crawls the web using a headless
4// browser controlled by the Playwright library.
5const crawler = new PlaywrightCrawler({
6    // Use the requestHandler to process each of the crawled pages.
7    async requestHandler({ request, page, enqueueLinks, log }) {
8        const title = await page.title();
9        log.info(`Title of ${request.loadedUrl} is '${title}'`);
10
11        // Save results as JSON to ./storage/datasets/default
12        await Dataset.pushData({ title, url: request.loadedUrl });
13
14        // Extract links from the current page
15        // and add them to the crawling queue.
16        await enqueueLinks();
17    },
18    // Uncomment this option to see the browser window.
19    // headless: false,
20});
21
22// Add first URL to the queue and start the crawl.
23await crawler.run(['https://crawlee.dev']);

By default, Crawlee stores data to ./storage in the current working directory. You can override this directory via Crawlee configuration. For details, see Configuration guide, Request storage and Result storage.

Installing pre-release versions

We provide automated beta builds for every merged code change in Crawlee. You can find them in the npm list of releases. If you want to test new features or bug fixes before we release them, feel free to install a beta build like this:

1npm install crawlee@3.12.3-beta.13

If you also use the Apify SDK, you need to specify dependency overrides in your package.json file so that you don't end up with multiple versions of Crawlee installed:

1{
2    "overrides": {
3       "apify": {
4           "@crawlee/core": "3.12.3-beta.13",
5           "@crawlee/types": "3.12.3-beta.13",
6           "@crawlee/utils": "3.12.3-beta.13"
7       }
8    }
9}

🛠 Features

Single interface for HTTP and headless browser crawling
Persistent queue for URLs to crawl (breadth & depth first)
Pluggable storage of both tabular data and files
Automatic scaling with available system resources
Integrated proxy rotation and session management
Lifecycles customizable with hooks
CLI to bootstrap your projects
Configurable routing, error handling and retries
Dockerfiles ready to deploy
Written in TypeScript with generics

👾 HTTP crawling

Zero config HTTP2 support, even for proxies
Automatic generation of browser-like headers
Replication of browser TLS fingerprints
Integrated fast HTML parsers. Cheerio and JSDOM
Yes, you can scrape JSON APIs as well

💻 Real browser crawling

JavaScript rendering and screenshots
Headless and headful support
Zero-config generation of human-like fingerprints
Automatic browser management
Use Playwright and Puppeteer with the same interface
Chrome, Firefox, Webkit and many others

Usage on the Apify platform

Crawlee is open-source and runs anywhere, but since it's developed by Apify, it's easy to set up on the Apify platform and run in the cloud. Visit the Apify SDK website to learn more about deploying Crawlee to the Apify platform.

Support

If you find any bug or issue with Crawlee, please submit an issue on GitHub. For questions, you can ask on Stack Overflow, in GitHub Discussions or you can join our Discord server.

Contributing

Your code contributions are welcome, and you'll be praised to eternity! If you have any ideas for improvements, either submit an issue or create a pull request. For contribution guidelines and the code of conduct, see CONTRIBUTING.md.

License

This project is licensed under the Apache License 2.0 - see the LICENSE.md file for details.

No vulnerabilities found.

No security vulnerabilities found.