• Fri frakt över 249 kr
  • •
  • Snabba leveranser
  • •
  • Billiga böcker
Kundservice

Du är på sajten för privatpersoner.

Företag, bibliotek eller offentlig verksamhet?

Du handlar på classic.bokus.com, där alla dina funktioner finns intakta.
Till classic.bokus.com
Bokus logotyp. Gå till startsidan.
  • Erbjudanden
  • Nyheter
  • Student
  • Topplistor
  • Barn & ungdom
  • Bokus Play
  • E-böcker
  • Pocketböcker
  • Spel & pussel

10% rabatt på allt med kod NYSTART10 →

Sidfot

Mina sidor

    Hjälp

    • Kundservice
    • Vanliga frågor och svar
    • Frakt och leverans
    • Retur vid ångerrätt
    • Reklamera vara
    • Betalning
    • Köpvillkor
    • Allmänna villkor
    • Information om webbplatsens tillgänglighet

    Om Bokus

    • Om oss
    • Pressrum
    • För studenter
    • För företag
    • För bibliotek och offentlig verksamhet
    • För leverantörer
    • Hållbarhet

    Populärt

    • Aktuella erbjudanden
    • Presentkort
    • Studentlitteratur
    • Nya böcker
    • Topplistor
    • Signerade böcker
    • Engelska böcker

    Inspiration

    • Boktips
    • BookTok
    • Populära bokserier
    • Barnbokskaraktärer
    • Populära författare
    Logotyp för Bokus
    Följ oss på Facebook (extern länk)Följ oss på Instagram (extern länk)Följ oss på YouTube (extern länk)Följ oss på TikTok (extern länk)
    bokus @ CookiesAnpassa cookiesIntegritetspolicyKöpvillkor
    Till Citymail hemsida (extern länk)Till Budbee hemsida (extern länk)Till Postnord hemsida (extern länk)Till Schenker hemsida (extern länk)Till Early Bird hemsida (extern länk)Till Walleys hemsida (extern länk)
    1. Data och IT
    2. Databaser

    Automated Data Collection with R

    A Practical Guide to Web Scraping and Text Mining

    AvSimon Munzert,Christian Rubba

    Inbunden, Engelska, 2014

    797 kr

    Beställningsvara. Skickas inom 5-8 vardagar. Fri frakt över 249 kr.

    Beskrivning

    A hands on guide to web scraping and text mining for both beginners and experienced users of R Introduces fundamental concepts of the main architecture of the web and databases and covers HTTP, HTML, XML, JSON, SQL.Provides basic techniques to query web documents and data sets (XPath and regular expressions).An extensive set of exercises are presented to guide the reader through each technique.Explores both supervised and unsupervised techniques as well as advanced techniques such as data scraping and text management.Case studies are featured throughout along with examples for each technique presented.R code and solutions to exercises featured in the book are provided on a supporting website.

    Produktinformation

    • Utgivningsdatum:2014-12-26
    • Mått:175 x 249 x 33 mm
    • Vikt:930 g
    • Format:Inbunden
    • Språk:Engelska
    • Antal sidor:480
    • Förlag:John Wiley & Sons Inc
    • ISBN:9781118834817

    Utforska kategorier

    • Databaser inom Data och IT

    Mer om författaren

    Simon Munzert is the author of Automated Data Collection with R: A Practical Guide to Web Scraping and Text Mining, published by Wiley.Christian Rubba is the author of Automated Data Collection with R: A Practical Guide to Web Scraping and Text Mining, published by Wiley.Peter Meißner is the author of Automated Data Collection with R: A Practical Guide to Web Scraping and Text Mining, published by Wiley.Dominic Nyhuis is the author of Automated Data Collection with R: A Practical Guide to Web Scraping and Text Mining, published by Wiley.

    Innehållsförteckning

    • Preface xv1 Introduction 11.1 Case study: World Heritage Sites in Danger 11.2 Some remarks on web data quality 71.3 Technologies for disseminating, extracting, and storing web data 91.4 Structure of the book 13Part One A Primer on Web and Data Technologies 152 HTML 172.1 Browser presentation and source code 182.2 Syntax rules 192.3 Tags and attributes 242.4 Parsing 323 XML and JSON 413.1 A short example XML document 423.2 XML syntax rules 433.3 When is an XML document well formed or valid? 513.4 XML extensions and technologies 533.5 XML and R in practice 603.6 A short example JSON document 683.7 JSON syntax rules 693.8 JSON and R in practice 714 XPath 794.1 XPath--a query language for web documents 804.2 Identifying node sets with XPath 814.3 Extracting node elements 935 HTTP 1015.1 HTTP fundamentals 1025.2 Advanced features of HTTP 1165.3 Protocols beyond HTTP 1245.4 HTTP in action 1266 AJAX 1496.1 JavaScript 1506.2 XHR 1546.3 Exploring AJAX with Web Developer Tools 1587 SQL and relational databases 1647.1 Overview and terminology 1657.2 Relational Databases 1677.3 SQL: a language to communicate with Databases 1757.4 Databases in action 1888 Regular expressions and essential string functions 1968.1 Regular expressions 1988.2 String processing 2078.3 A word on character encodings 214Part Two A Practical Toolbox forWeb Scraping and Text Mining 2199 Scraping the Web 2219.1 Retrieval scenarios 2229.2 Extraction strategies 2709.3 Web scraping: Good practice 2789.4 Valuable sources of inspiration 29010 Statistical text processing 29510.1 The running example: Classifying press releases of the British government 29610.2 Processing textual data 29810.3 Supervised learning techniques 30710.4 Unsupervised learning techniques 31311 Managing data projects 32211.1 Interacting with the file system 32211.2 Processing multiple documents/links 32311.3 Organizing scraping procedures 32811.4 Executing R scripts on a regular basis 334Part Three A Bag of Case Studies 34112 Collaboration networks in the US Senate 34312.1 Information on the bills 34412.2 Information on the senators 35012.3 Analyzing the network structure 35312.4 Conclusion 35813 Parsing information from semistructured documents 35913.1 Downloading data from the FTP server 36013.2 Parsing semistructured text data 36113.3 Visualizing station and temperature data 36814 Predicting the 2014 Academy Awards using Twitter 37115 Mapping the geographic distribution of names 38015.1 Developing a data collection strategy 38115.2 Website inspection 38215.3 Data retrieval and information extraction 38415.4 Mapping names 38715.5 Automating the process 38916 Gathering data on mobile phones 39616.1 Page exploration 39616.2 Scraping procedure 40416.3 Graphical analysis 40616.4 Data storage 40817 Analyzing sentiments of product reviews 41617.1 Introduction 41617.2 Collecting the data 41717.3 Analyzing the data 42617.4 Conclusion 434References 435General index 442Package index 448Function index 449