Giter VIP home page Giter VIP logo

crawl-anywhere's Introduction

Crawl-Anywhere

April 2013 - Starting version 4.0, Crawl-Anywhere becomes an open-source project. Current version is 4.0.0 release candidate

Stable version 3.0.x is still available at http://www.crawl-anywhere.com/

Introduction

Crawl Anywhere is mainly a web crawler. However, Crawl-Anywhere includes all components in order to build a vertical search engine.

Crawl Anywhere includes :

Project home page : http://www.crawl-anywhere.com/

A web crawler is a program that discovers and read all HTML pages or documents (HTML, PDF, Office, ...) on a web site in order for example to index these data and build a search engine (like google). Wikipedia provides a great description of what is a Web crawler : http://en.wikipedia.org/wiki/Web_crawler.

Support

Build distribution

Pre-requisites :

  • Maven 3.0.0 or >
  • Oracle Java 6 or >

Steps :

Installation

Pre-requisites :

  • Oracle Java 6 or >
  • Tomcat 5.5 or >
  • Apache 2.0 or >
  • PHP 5.2.x or 5.3.x or 5.4.x
  • MongoDB 64 bits 2.2 or >
  • Solr 3.x or > (configuration files provided for Solr 4.3.0)

Steps :

Getting Started

See the User Manual at http://www.crawl-anywhere.com/getting-started/

History

  • release 4.0.0-alpha-1 : April, 28 2013
  • release 4.0.0-alpha-2 : May, 22 2013
  • release 4.0.0-alpha-3 : June, 21 2013
  • release 4.0.0-alpha-4 : June, 23 2013
  • release 4.0.0-beta-1 : August, 6 2013
  • release 4.0.0-release-candidate : October, 20 2013

Roadmap

  • release 4.0.0 final : December 2013

crawl-anywhere's People

Contributors

bejean avatar

Watchers

James Cloos avatar  avatar

Recommend Projects

  • React photo React

    A declarative, efficient, and flexible JavaScript library for building user interfaces.

  • Vue.js photo Vue.js

    ๐Ÿ–– Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.

  • Typescript photo Typescript

    TypeScript is a superset of JavaScript that compiles to clean JavaScript output.

  • TensorFlow photo TensorFlow

    An Open Source Machine Learning Framework for Everyone

  • Django photo Django

    The Web framework for perfectionists with deadlines.

  • D3 photo D3

    Bring data to life with SVG, Canvas and HTML. ๐Ÿ“Š๐Ÿ“ˆ๐ŸŽ‰

Recommend Topics

  • javascript

    JavaScript (JS) is a lightweight interpreted programming language with first-class functions.

  • web

    Some thing interesting about web. New door for the world.

  • server

    A server is a program made to process requests and deliver data to clients.

  • Machine learning

    Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.

  • Game

    Some thing interesting about game, make everyone happy.

Recommend Org

  • Facebook photo Facebook

    We are working to build community through open source technology. NB: members must have two-factor auth.

  • Microsoft photo Microsoft

    Open source projects and samples from Microsoft.

  • Google photo Google

    Google โค๏ธ Open Source for everyone.

  • D3 photo D3

    Data-Driven Documents codes.