- 专利标题: Optimizing web crawling through web page pruning
-
申请号: US15688167申请日: 2017-08-28
-
公开(公告)号: US09996619B2公开(公告)日: 2018-06-12
- 发明人: Shahar Sperling , Omer Tripp , Omri Weisman
- 申请人: International Business Machines Corporation
- 申请人地址: US NY Armonk
- 专利权人: INTERNATIONAL BUSINESS MACHINES CORPORATION
- 当前专利权人: INTERNATIONAL BUSINESS MACHINES CORPORATION
- 当前专利权人地址: US NY Armonk
- 代理机构: Cantor Colburn LLP
- 代理商 Maeve Carpenter
- 主分类号: G06F17/30
- IPC分类号: G06F17/30 ; G06F17/22
摘要:
Crawling computer-based documents by performing static analysis on a computer-based document to identify within the computer-based document one or more execution vectors, where each execution vector includes a computer program segment including a call to an entity that is external to the computer-based document, and one or more additional computer program segments whose execution precedes and leads ultimately to execution of the computer program segment that includes the call to the entity, and causing any of the computer program segments in any of the execution vectors to be executed during a crawling of the computer-based document, and any computer program segment within the computer-based document that is excluded from the execution vectors to be excluded from execution during the crawling of the computer-based document.
公开/授权文献
- US20170351761A1 OPTIMIZING WEB CRAWLING THROUGH WEB PAGE PRUNING 公开/授权日:2017-12-07
信息查询