Tài liệuTrình duyệt và CrawlAgent BrowserTối ưu hóa chi phí

Tối ưu Chi phí

Giới thiệu

Khi sử dụng Puppeteer để thu thập dữ liệu, mức tiêu thụ lưu lượng là một yếu tố quan trọng cần cân nhắc. Đặc biệt khi sử dụng dịch vụ proxy, chi phí lưu lượng có thể tăng lên đáng kể. Để tối ưu việc sử dụng lưu lượng, chúng ta có thể áp dụng các chiến lược sau:

  1. Chặn tài nguyên: Giảm mức tiêu thụ lưu lượng bằng cách chặn các yêu cầu tài nguyên không cần thiết.
  2. Chặn URL yêu cầu: Giảm thêm lưu lượng bằng cách chặn các yêu cầu cụ thể dựa trên đặc điểm của URL.
  3. Mô phỏng thiết bị di động: Sử dụng cấu hình thiết bị di động để lấy các phiên bản trang nhẹ hơn.
  4. Tối ưu tổng hợp: Kết hợp các phương pháp trên để đạt kết quả tốt nhất.

Giải pháp Tối ưu 1: Chặn Tài nguyên

Giới thiệu về Chặn Tài nguyên

Trong Puppeteer, page.setRequestInterception(true) có thể bắt mọi yêu cầu mạng do trình duyệt khởi tạo và quyết định tiếp tục (request.continue()), hủy (request.abort()), hoặc tùy chỉnh phản hồi (request.respond()).

Phương pháp này có thể giảm đáng kể mức tiêu thụ băng thông, đặc biệt phù hợp cho các tình huống thu thập dữ liệu (crawling), chụp màn hình (screenshotting), và tối ưu hiệu năng.

Các Loại Tài nguyên Có thể Chặn và Đề xuất

Loại Tài nguyênMô tảVí dụTác động Sau khi ChặnĐề xuất
imageTài nguyên hình ảnhHình ảnh JPG/PNG/GIF/WebPHình ảnh sẽ không được hiển thị⭐ An toàn
fontTệp phông chữPhông chữ TTF/WOFF/WOFF2Sẽ dùng phông chữ mặc định của hệ thống thay thế⭐ An toàn
mediaTệp phương tiệnTệp video/âm thanhNội dung phương tiện không thể phát⭐ An toàn
manifestWeb App ManifestTệp cấu hình PWAChức năng PWA có thể bị ảnh hưởng⭐ An toàn
prefetchTài nguyên nạp trước<link rel="prefetch">Tác động tối thiểu đến trang⭐ An toàn
stylesheetBảng định kiểu CSSTệp CSS bên ngoàiKiểu dáng trang bị mất, có thể ảnh hưởng đến bố cục⚠️ Thận trọng
websocketWebSocketKết nối giao tiếp thời gian thựcChức năng thời gian thực bị vô hiệu hóa⚠️ Thận trọng
eventsourceServer-Sent EventsDữ liệu đẩy từ máy chủChức năng đẩy bị vô hiệu hóa⚠️ Thận trọng
preflightYêu cầu preflight CORSYêu cầu OPTIONSYêu cầu cross-origin thất bại⚠️ Thận trọng
scriptScript JavaScriptTệp JS bên ngoàiChức năng động bị vô hiệu hóa, SPA có thể không render được❌ Tránh
xhrYêu cầu XHRYêu cầu dữ liệu AJAXKhông thể lấy được dữ liệu động❌ Tránh
fetchYêu cầu FetchYêu cầu AJAX hiện đạiKhông thể lấy được dữ liệu động❌ Tránh
documentTài liệu chínhBản thân trang HTMLTrang không thể tải❌ Tránh

Giải thích Mức độ Đề xuất:

  • ⭐ An toàn: Việc chặn hầu như không ảnh hưởng đến việc thu thập dữ liệu hoặc render màn hình đầu tiên; nên chặn theo mặc định.
  • ⚠️ Thận trọng: Có thể phá vỡ kiểu dáng, chức năng thời gian thực, hoặc yêu cầu cross-origin; cần cân nhắc theo nghiệp vụ.
  • ❌ Tránh: Khả năng cao khiến các trang SPA/động không thể render hoặc lấy dữ liệu bình thường, trừ khi bạn hoàn toàn chắc chắn rằng không cần các tài nguyên này.

Mã Ví dụ Chặn Tài nguyên

import { Scrapeless } from '@scrapeless-ai/sdk';
import puppeteer from 'puppeteer-core';
 
const client = new Scrapeless({ apiKey: 'API Key' });
 
const { browserWSEndpoint } = client.browser.create({
    sessionName: 'sdk_test',
    sessionTTL: 180,
    proxyCountry: 'ANY',
    sessionRecording: true,
    fingerprint,
});
 
async function scrapeWithResourceBlocking(url) {
    const browser = await puppeteer.connect({
        browserWSEndpoint,
        defaultViewport: null
    });
    const page = await browser.newPage();
 
    // Enable request interception
    await page.setRequestInterception(true);
 
    // Define resource types to block
    const BLOCKED_TYPES = new Set([
        'image',
        'font',
        'media',
        'stylesheet',
    ]);
 
    // Intercept requests
    page.on('request', (request) => {
        if (BLOCKED_TYPES.has(request.resourceType())) {
            request.abort();
            console.log(`Blocked: ${request.resourceType()} - ${request.url().substring(0, 50)}...`);
        } else {
            request.continue();
        }
    });
 
    await page.goto(url, {waitUntil: 'domcontentloaded'});
 
    // Extract data
    const data = await page.evaluate(() => {
        return {
            title: document.title,
            content: document.body.innerText.substring(0, 1000)
        };
    });
 
    await browser.close();
    return data;
}
 
// Usage
scrapeWithResourceBlocking('https://www.scrapeless.com')
    .then(data => console.log('Scraping result:', data))
    .catch(error => console.error('Scraping failed:', error));

Giải pháp Tối ưu 2: Chặn URL Yêu cầu

Ngoài việc chặn theo loại tài nguyên, bạn có thể thực hiện kiểm soát chặn chi tiết hơn dựa trên đặc điểm của URL. Cách này đặc biệt hiệu quả để chặn quảng cáo, script phân tích, và các yêu cầu bên thứ ba không cần thiết khác.

Chiến lược Chặn URL

  1. Chặn theo tên miền: Chặn tất cả yêu cầu từ một tên miền cụ thể
  2. Chặn theo đường dẫn: Chặn các yêu cầu từ một đường dẫn cụ thể
  3. Chặn theo loại tệp: Chặn các tệp có phần mở rộng cụ thể
  4. Chặn theo từ khóa: Chặn các yêu cầu có URL chứa từ khóa cụ thể

Các Mẫu URL Thường có thể Chặn

Mẫu URLMô tảVí dụĐề xuất
Dịch vụ quảng cáoTên miền mạng quảng cáoad.doubleclick.net, googleadservices.com⭐ An toàn
Dịch vụ phân tíchScript thống kê và phân tíchgoogle-analytics.com, hotjar.com⭐ An toàn
Plugin mạng xã hộiNút chia sẻ mạng xã hội, v.v.platform.twitter.com, connect.facebook.net⭐ An toàn
Pixel theo dõiPixel theo dõi hành vi người dùngURL chứa pixel, beacon, tracker⭐ An toàn
Tệp phương tiện lớnTệp video, âm thanh lớnPhần mở rộng như .mp4, .webm, .mp3⭐ An toàn
Dịch vụ phông chữDịch vụ phông chữ trực tuyếnfonts.googleapis.com, use.typekit.net⭐ An toàn
Tài nguyên CDNCDN tài nguyên tĩnhcdn.jsdelivr.net, unpkg.com⚠️ Thận trọng

Mã Ví dụ Chặn URL

import puppeteer from 'puppeteer-core';
import { Scrapeless } from '@scrapeless-ai/sdk';
import puppeteer from 'puppeteer-core';
 
const client = new Scrapeless({ apiKey: 'API Key' });
 
const { browserWSEndpoint } = client.browser.create({
    sessionName: 'sdk_test',
    sessionTTL: 180,
    proxyCountry: 'ANY',
    sessionRecording: true,
    fingerprint,
});
 
async function scrapeWithUrlBlocking(url) {
    const browser = await puppeteer.connect({
        browserWSEndpoint,
        defaultViewport: null
    });
    const page = await browser.newPage();
 
    // Enable request interception
    await page.setRequestInterception(true);
 
    // Define domains and URL patterns to block
    const BLOCKED_DOMAINS = [
        'google-analytics.com',
        'googletagmanager.com',
        'doubleclick.net',
        'facebook.net',
        'twitter.com',
        'linkedin.com',
        'adservice.google.com',
    ];
 
    const BLOCKED_PATHS = [
        '/ads/',
        '/analytics/',
        '/pixel/',
        '/tracking/',
        '/stats/',
    ];
 
    // Intercept requests
    page.on('request', (request) => {
        const url = request.url();
 
        // Check domain
        if (BLOCKED_DOMAINS.some(domain => url.includes(domain))) {
            request.abort();
            console.log(`Blocked domain: ${url.substring(0, 50)}...`);
            return;
        }
 
        // Check path
        if (BLOCKED_PATHS.some(path => url.includes(path))) {
            request.abort();
            console.log(`Blocked path: ${url.substring(0, 50)}...`);
            return;
        }
 
        // Allow other requests
        request.continue();
    });
 
    await page.goto(url, {waitUntil: 'domcontentloaded'});
 
    // Extract data
    const data = await page.evaluate(() => {
        return {
            title: document.title,
            content: document.body.innerText.substring(0, 1000)
        };
    });
 
    await browser.close();
    return data;
}
 
// Usage
scrapeWithUrlBlocking('https://www.scrapeless.com')
    .then(data => console.log('Scraping result:', data))
    .catch(error => console.error('Scraping failed:', error));

Giải pháp Tối ưu 3: Mô phỏng Thiết bị Di động

Mô phỏng thiết bị di động là một chiến lược tối ưu lưu lượng hiệu quả khác vì các trang web dành cho di động thường cung cấp nội dung trang nhẹ hơn.

Ưu điểm của Mô phỏng Thiết bị Di động

  1. Phiên bản trang nhẹ hơn: Nhiều trang web cung cấp nội dung ngắn gọn hơn cho thiết bị di động
  2. Tài nguyên hình ảnh nhỏ hơn: Phiên bản di động thường tải các hình ảnh nhỏ hơn
  3. CSS và JavaScript được đơn giản hóa: Phiên bản di động thường dùng kiểu dáng và script được đơn giản hóa
  4. Giảm quảng cáo và nội dung không cốt lõi: Phiên bản di động thường loại bỏ một số chức năng không cốt lõi
  5. Phản hồi thích ứng: Lấy bố cục nội dung được tối ưu cho màn hình nhỏ

Cấu hình Mô phỏng Thiết bị Di động

Dưới đây là các tham số cấu hình cho một số thiết bị di động thường dùng:

const iPhoneX = {
    viewport: {
        width: 375,
        height: 812,
        deviceScaleFactor: 3,
        isMobile: true,
        hasTouch: true,
        isLandscape: false
    }
};

Hoặc trực tiếp sử dụng các phương thức tích hợp sẵn của puppeteer để mô phỏng thiết bị di động

import { KnownDevices } from 'puppeteer-core';
const iPhone = KnownDevices['iPhone 15 Pro'];
 
const browser = await puppeteer.launch();
const page = await browser.newPage();
await page.emulate(iPhone);

Mã Ví dụ Mô phỏng Thiết bị Di động

import puppeteer from 'puppeteer-core';
import { Scrapeless } from '@scrapeless-ai/sdk';
 
const client = new Scrapeless({ apiKey: 'API Key' });
 
const { browserWSEndpoint } = client.browser.create({
    sessionName: 'sdk_test',
    sessionTTL: 180,
    proxyCountry: 'ANY',
    sessionRecording: true,
    fingerprint,
});
 
async function scrapeWithMobileEmulation(url) {
    const browser = await puppeteer.connect({
        browserWSEndpoint,
        defaultViewport: null
    });
 
    const page = await browser.newPage();
 
    // Set mobile device simulation
    const iPhone = KnownDevices['iPhone 15 Pro'];
    await page.emulate(iPhone);
 
    await page.goto(url, {waitUntil: 'domcontentloaded'});
    // Extract data
    const data = await page.evaluate(() => {
        return {
            title: document.title,
            content: document.body.innerText.substring(0, 1000)
        };
    });
 
    await browser.close();
    return data;
}
 
// Usage
scrapeWithMobileEmulation('https://www.scrapeless.com')
    .then(data => console.log('Scraping result:', data))
    .catch(error => console.error('Scraping failed:', error));

Ví dụ Tối ưu Tổng hợp

Dưới đây là một ví dụ tổng hợp kết hợp tất cả các giải pháp tối ưu:

import puppeteer, {KnownDevices} from 'puppeteer-core';
import { Scrapeless } from '@scrapeless-ai/sdk';
 
const client = new Scrapeless({ apiKey: 'API Key' });
 
const { browserWSEndpoint } = client.browser.create({
    sessionName: 'sdk_test',
    sessionTTL: 180,
    proxyCountry: 'ANY',
    sessionRecording: true,
    fingerprint,
});
 
async function optimizedScraping(url) {
    console.log(`Starting optimized scraping: ${url}`);
 
    // Record traffic usage
    let totalBytesUsed = 0;
 
    const browser = await puppeteer.connect({
        browserWSEndpoint,
        defaultViewport: null
    });
 
    const page = await browser.newPage();
 
    // Set mobile device simulation
    const iPhone = KnownDevices['iPhone 15 Pro'];
    await page.emulate(iPhone);
 
    // Set request interception
    await page.setRequestInterception(true);
 
    // Define resource types to block
    const BLOCKED_TYPES = [
        'image',
        'media',
        'font'
    ];
 
    // Define domains to block
    const BLOCKED_DOMAINS = [
        'google-analytics.com',
        'googletagmanager.com',
        'facebook.net',
        'doubleclick.net',
        'adservice.google.com'
    ];
 
    // Define URL paths to block
    const BLOCKED_PATHS = [
        '/ads/',
        '/analytics/',
        '/tracking/'
    ];
 
    // Intercept requests
    page.on('request', (request) => {
        const url = request.url();
        const resourceType = request.resourceType();
 
        // Check resource type
        if (BLOCKED_TYPES.includes(resourceType)) {
            console.log(`Blocked resource type: ${resourceType} - ${url.substring(0, 50)}...`);
            request.abort();
            return;
        }
 
        // Check domain
        if (BLOCKED_DOMAINS.some(domain => url.includes(domain))) {
            console.log(`Blocked domain: ${url.substring(0, 50)}...`);
            request.abort();
            return;
        }
 
        // Check path
        if (BLOCKED_PATHS.some(path => url.includes(path))) {
            console.log(`Blocked path: ${url.substring(0, 50)}...`);
            request.abort();
            return;
        }
 
        // Allow other requests
        request.continue();
    });
 
    // Monitor network traffic
    page.on('response', async (response) => {
        const headers = response.headers();
        const contentLength = headers['content-length'] ? parseInt(headers['content-length'], 10) : 0;
        totalBytesUsed += contentLength;
    });
 
    await page.goto(url, {waitUntil: 'domcontentloaded'});
 
    // Simulate scrolling to trigger lazy-loading content
    await page.evaluate(() => {
        window.scrollBy(0, window.innerHeight);
    });
 
    await new Promise(resolve => setTimeout(resolve, 1000))
 
    // Extract data
    const data = await page.evaluate(() => {
        return {
            title: document.title,
            content: document.body.innerText.substring(0, 1000),
            links: Array.from(document.querySelectorAll('a')).slice(0, 10).map(a => ({
                text: a.innerText,
                href: a.href
            }))
        };
    });
 
    // Output traffic usage statistics
    console.log(`\nTraffic Usage Statistics:`);
    console.log(`Used: ${(totalBytesUsed / 1024 / 1024).toFixed(2)} MB`);
 
    await browser.close();
    return data;
}
 
// Usage
optimizedScraping('https://www.scrapeless.com')
    .then(data => console.log('Scraping complete:', data))
    .catch(error => console.error('Scraping failed:', error));

So sánh Tối ưu

Chúng ta thử loại bỏ mã tối ưu khỏi ví dụ tổng hợp để so sánh lưu lượng trước và sau khi tối ưu. Dưới đây là mã ví dụ chưa được tối ưu:

import puppeteer from 'puppeteer-core';
import { Scrapeless } from '@scrapeless-ai/sdk';
 
const client = new Scrapeless({ apiKey: 'API Key' });
 
const { browserWSEndpoint } = client.browser.create({
    sessionName: 'sdk_test',
    sessionTTL: 180,
    proxyCountry: 'ANY',
    sessionRecording: true,
    fingerprint,
});
 
async function unoptimizedScraping(url) {
  console.log(`Starting unoptimized scraping: ${url}`);
 
  // Record traffic usage
  let totalBytesUsed = 0;
 
  const browser = await puppeteer.connect({
    browserWSEndpoint,
    defaultViewport: null
  });
 
  const page = await browser.newPage();
 
  // Set request interception
  await page.setRequestInterception(true);
 
  // Intercept requests
  page.on('request', (request) => {
    request.continue();
  });
 
  // Monitor network traffic
  page.on('response', async (response) => {
    const headers = response.headers();
    const contentLength = headers['content-length'] ? parseInt(headers['content-length'], 10) : 0;
    totalBytesUsed += contentLength;
  });
 
  await page.goto(url, {waitUntil: 'domcontentloaded'});
 
  // Simulate scrolling to trigger lazy-loading content
  await page.evaluate(() => {
    window.scrollBy(0, window.innerHeight);
  });
 
  await new Promise(resolve => setTimeout(resolve, 1000))
 
  // Extract data
  const data = await page.evaluate(() => {
    return {
      title: document.title,
      content: document.body.innerText.substring(0, 1000),
      links: Array.from(document.querySelectorAll('a')).slice(0, 10).map(a => ({
        text: a.innerText,
        href: a.href
      }))
    };
  });
 
  // Output traffic usage statistics
  console.log(`\nTraffic Usage Statistics:`);
  console.log(`Used: ${(totalBytesUsed / 1024 / 1024).toFixed(2)} MB`);
 
  await browser.close();
  return data;
}
 
// Usage
unoptimizedScraping('https://www.scrapeless.com')
  .then(data => console.log('Scraping complete:', data))
  .catch(error => console.error('Scraping failed:', error));

Sau khi chạy mã chưa được tối ưu, chúng ta có thể thấy rất trực quan sự khác biệt về lưu lượng từ thông tin được in ra:

Tình huốngLưu lượng Đã dùng (MB)Tỷ lệ Tiết kiệm
Chưa tối ưu6.03—
Đã tối ưu0.81≈ 86.6 %

Bằng cách kết hợp các giải pháp tối ưu trên, mức tiêu thụ lưu lượng proxy có thể được giảm đáng kể, hiệu quả thu thập dữ liệu có thể được cải thiện, đồng thời đảm bảo lấy được nội dung cốt lõi cần thiết.