Node.js SDK

Tổng quan

Scrapeless Node.js SDK chính thức cung cấp quyền truy cập vào tự động hóa trình duyệt, thu thập dữ liệu, crawl, proxy, kết quả tìm kiếm và trích xuất nội dung AI chat. Hướng dẫn này bám theo README của kho lưu trữ SDK, với các ví dụ có thể chạy được và chi tiết cấu hình.

Yêu cầu

Sử dụng Node.js với npm, pnpm hoặc Yarn. Package manifest không khai báo phiên bản Node.js tối thiểu; quy trình phát hành của kho lưu trữ sử dụng Node.js 20. JavaScript và TypeScript đều được hỗ trợ, với cả bản xuất ES module và CommonJS.

Các ví dụ bên dưới sử dụng ES module. Hãy lưu chúng dưới dạng tệp .mjs, hoặc đặt "type": "module" trong tệp package.json của dự án.

Cài đặt

npm install @scrapeless-ai/sdk

Bạn cũng có thể dùng pnpm add @scrapeless-ai/sdk hoặc yarn add @scrapeless-ai/sdk.

Xác thực / API Key

Đăng nhập vào bảng điều khiển Scrapeless và tạo một API key. Hãy export nó trước khi chạy các ví dụ:

export SCRAPELESS_API_KEY="YOUR_API_KEY"

Hãy giữ API key của bạn trong biến môi trường hoặc trình quản lý bí mật thay vì commit nó vào hệ thống quản lý mã nguồn.

new Scrapeless() đọc SCRAPELESS_API_KEY. Bạn cũng có thể truyền tùy chọn apiKey khi khởi tạo client.

Bắt đầu nhanh

Lưu tệp này dưới dạng quickstart.mjs và chạy node quickstart.mjs sau khi đã thiết lập API key của bạn.

import { Scrapeless } from '@scrapeless-ai/sdk';
 
const client = new Scrapeless();
const result = await client.universal.scrape({
  actor: 'unlocker.webunlocker',
  input: { url: 'https://example.com', method: 'GET', redirect: false }
});
console.log(result);

Ma trận phạm vi sản phẩm

Sản phẩmDịch vụ SDKPhạm vi
Agent Browserclient.browserTạo và quản lý các phiên trình duyệt từ xa.
Browser Profilesclient.profilesLưu giữ dữ liệu trình duyệt qua các phiên.
Scraping APIclient.scrapingTrích xuất dữ liệu có cấu trúc bằng các actor cho website.
Web Unlockerclient.universalTruy xuất nội dung từ các website được bảo vệ.
Crawlclient.scrapingCrawlThu thập dữ liệu một trang hoặc crawl toàn bộ website.
Google Search APIclient.deepserpTrích xuất kết quả từ công cụ tìm kiếm.
Proxiesclient.proxiesTạo các URL kết nối proxy.
AI Scraperclient.aiScraperTạo các tác vụ AI chat và truy xuất trạng thái cùng kết quả của chúng.

Ví dụ sử dụng

Trừ khi một ví dụ tự khởi tạo client riêng, hãy tái sử dụng const client = new Scrapeless() từ phần Bắt đầu nhanh.

Browser

Cài đặt Puppeteer cho ví dụ này: npm install puppeteer-core. Đối với Playwright và các trình bao bọc trình duyệt của SDK, xem các ví dụ tích hợp trình duyệt.

Quản lý phiên trình duyệt nâng cao hỗ trợ các framework Playwright và Puppeteer, với các khả năng chống phát hiện có thể cấu hình (ví dụ: giả mạo fingerprint, giải CAPTCHA) và các quy trình tự động hóa có thể mở rộng:

import { Scrapeless } from '@scrapeless-ai/sdk';
import puppeteer from 'puppeteer-core';
 
const client = new Scrapeless();
 
// Create a browser session
const { browserWSEndpoint } = await client.browser.create({
  sessionName: 'my-session',
  sessionTTL: 180,
  proxyCountry: 'US'
});
 
// Connect with Puppeteer
const browser = await puppeteer.connect({
  browserWSEndpoint: browserWSEndpoint
});
 
try {
  const page = await browser.newPage();
  await page.goto('https://example.com');
  console.log(await page.title());
} finally {
  await browser.close();
}

Browser Profile

Quản lý các browser profile cho các phiên bền vững.

const createResponse = await client.profiles.create('My Profile');
console.log('Profile created:', createResponse);
 
const profiles = await client.profiles.list({ page: 1, pageSize: 10 });
console.log('Profiles:', profiles.docs);
 
const profile = await client.profiles.get(createResponse.profileId);
console.log('Profile details:', profile);
 
// Delete the profile when it is no longer needed.
await client.profiles.delete(createResponse.profileId);

Scraping API

Các API trích xuất dữ liệu trực tiếp cho website (ví dụ: thương mại điện tử, nền tảng du lịch). Truy xuất thông tin sản phẩm có cấu trúc, giá cả và đánh giá với các trình kết nối được dựng sẵn:

const result = await client.scraping.scrape({
  actor: 'scraper.google.search',
  input: {
    'q': 'coffee',
    'hl': 'en',
    'gl': 'us'
  }
});
 
console.log(result.data);

Web Unlocker

Trích xuất dữ liệu từ các website bằng Web Unlocker (được cung cấp dưới dạng client.universal).

const result = await client.universal.scrape({
  actor: 'unlocker.webunlocker',
  input: { url: 'https://example.com', method: 'GET', redirect: false }
});
console.log(result);

Crawl

Trích xuất dữ liệu từ từng trang riêng lẻ hoặc duyệt toàn bộ tên miền, xuất ra ở các định dạng bao gồm Markdown, JSON, HTML, ảnh chụp màn hình và liên kết.

const result = await client.scrapingCrawl.scrapeUrl('https://example.com');
 
console.log(result);

Proxy

Tạo một URL proxy bằng cách sử dụng thiết lập gateway và phiên của bạn.

const proxyUrl = client.proxies.proxy({
  type: 'residential',
  country: 'US',
  sessionDuration: 30,
  sessionId: client.proxies.generateSessionId(),
  gateway: 'your-proxy-gateway:port'
});
console.log(proxyUrl);

AI Scraper

Tạo một tác vụ: client.aiScraper.createTask(request)

Truyền vào actor bắt buộc và input đặc thù cho actor. Một đối tượng webhook tùy chọn chấp nhận một url callback. Promise phân giải thành phản hồi API đầy đủ, bao gồm task_id và status, cùng task_result khi khả dụng.

Trích xuất nội dung AI chat hàng loạt để theo dõi các lượt nhắc đến thương hiệu, so sánh câu trả lời và phân tích thông tin cạnh tranh từ các mô hình mới nhất. Truy xuất URL, prompt, câu trả lời dạng Markdown, trích dẫn và nhiều hơn nữa thông qua một tích hợp duy nhất.

Các actor được hỗ trợ bao gồm scraper.chatgpt, scraper.perplexity, scraper.copilot, scraper.gemini, scraper.aimode, scraper.overview, scraper.grok và scraper.alexa. JSON input phụ thuộc vào actor; xem tài liệu AI Scraper để biết chi tiết các tham số. JSON webhook tùy chọn chứa một url callback.

import { Scrapeless } from '@scrapeless-ai/sdk';
 
const client = new Scrapeless(); // Uses SCRAPELESS_API_KEY
 
const task = await client.aiScraper.createTask({
  actor: 'scraper.chatgpt',
  input: {
    prompt: 'Most reliable proxy service for data extraction',
    country: 'US',
    web_search: true
  },
  // Optional: webhook: { url: 'https://your-webhook.example.com' }
});
console.log('Created task:', task);

Lấy trạng thái và kết quả của tác vụ: client.aiScraper.getTaskResult(taskId)

Truyền vào task_id từ bước tạo. Tiếp tục trong cùng một script, hoặc lưu ID và truy xuất kết quả trong một yêu cầu sau đó.

const result = await client.aiScraper.getTaskResult(task.task_id);
 
switch (result.status) {
  case 'success':
    console.log('Task result:', result.task_result);
    break;
  case 'failed':
    console.error('Task failed:', result.message);
    break;
  case 'running':
    console.log('Task is running. Retrieve the result again later.');
    break;
}

Cả hai phương thức đều trả về JSON API không thay đổi. Bước tạo trả về task_id, status, và, khi khả dụng, task_result. Việc truy xuất kết quả trả về status, task_result khi khả dụng, và message khi thất bại. Trạng thái là success, failed hoặc running; SDK không tự động polling.

Trạng tháiÝ nghĩaBước tiếp theo
runningTác vụ vẫn đang được xử lý.Gọi lại getTaskResult sau đó hoặc sử dụng webhook.
successTác vụ đã hoàn tất.Đọc task_result; cấu trúc của nó phụ thuộc vào actor.
failedTác vụ không thể hoàn tất.Đọc message để biết lý do thất bại.

Bước tạo có thể đã bao gồm sẵn một kết quả. Hãy kiểm tra trạng thái của nó trước khi lên lịch các yêu cầu tiếp theo. Nếu bạn triển khai polling, hãy sử dụng độ trễ và một thời gian chờ tổng thể.

Google Search API

const result = await client.deepserp.scrape({
  actor: 'scraper.google.search',
  input: { q: 'nike site:www.nike.com' }
});
console.log(result);

Để có các tích hợp hoàn chỉnh hơn, hãy duyệt thư mục examples của kho lưu trữ.

Xử lý lỗi

Bắt ScrapelessError cho các lỗi yêu cầu API. Hãy kiểm tra status của phản hồi AI Scraper một cách riêng biệt: một tác vụ có thể trả về failed mà yêu cầu HTTP không ném ra lỗi.

import { Scrapeless, ScrapelessError } from '@scrapeless-ai/sdk';
 
try {
  const client = new Scrapeless();
  const result = await client.universal.scrape({
    actor: 'unlocker.webunlocker',
    input: { url: 'https://example.com', method: 'GET' }
  });
  console.log(result);
} catch (error) {
  if (error instanceof ScrapelessError) {
    console.error('Scrapeless error:', error.message);
    console.error('Status code:', error.statusCode);
  } else {
    throw error;
  }
}

Cấu hình / Biến môi trường

API key là bắt buộc. Việc ghi đè endpoint là tùy chọn; bảng dưới đây hiển thị các giá trị mặc định của chúng.

import { Scrapeless } from '@scrapeless-ai/sdk';
 
const client = new Scrapeless({
  apiKey: process.env.SCRAPELESS_API_KEY,
  timeout: 30000, // Request timeout in milliseconds
  baseApiUrl: 'https://api.scrapeless.com',
  browserApiUrl: 'https://browser.scrapeless.com',
  scrapingCrawlApiUrl: 'https://api.scrapeless.com'
});

Cấu hình được chỉ định tường minh có độ ưu tiên cao hơn biến môi trường. Thời gian chờ yêu cầu mặc định là 30.000 mili giây.

Biến môi trườngMục đích / mặc định
SCRAPELESS_API_KEYAPI key bắt buộc từ bảng điều khiển.
SCRAPELESS_BASE_API_URLhttps://api.scrapeless.com
SCRAPELESS_BROWSER_API_URLhttps://browser.scrapeless.com
SCRAPELESS_CRAWL_API_URLhttps://api.scrapeless.com

Hỗ trợ

SDK được phát hành theo Giấy phép MIT.

Các dự án liên quan