Showing posts with label information retrieval. Show all posts
Showing posts with label information retrieval. Show all posts

An Introduction to Search Engines and Web Navigation Review

An Introduction to Search Engines and Web Navigation
Average Reviews:

(More customer reviews)
The first chapters of the book provide an extensive introduction to search engines and navigation. No formal prerequisites are required; any Web enthusiast will enjoy reading the book. These chapters comprise of background and history of the Web, navigation and searching, search engine architecture and different types of search engines. In addition to the basics, additional topics covered are navigation (aka surfing), the interplay between search and navigation, Web data mining, personalization, the mobile web, social networks, collaborative filtering and Weblogs (aka Blogs).
The book goes far beyond simple searching and navigation; it provides a comprehensive overview of the current research fronts in areas related to Web search engines and navigation.
The text is highly readable with a large number of illustrations and examples. It can serve as an excellent textbook both for an introductory and a more advanced course of Web search and navigation. Each chapter starts with a listing of objectives and ends with a set of exercises relevant to the topics covered in the chapter. Students will especially benefit from the non-technical descriptions and clear explanations of the concepts.
The book is also a great reference source for researchers and IT professionals: it includes 410 references to articles, and 202 references to Web pages and resources. I highly recommend the book.


Click Here to see more reviews about: An Introduction to Search Engines and Web Navigation

This book is a second edition, updated and expanded toexplain the technologies that help us find information on the web. Search engines and web navigation tools have become ubiquitous in our day to day use of the web as an information source, a tool for commercial transactions and a social computing tool. Moreover, through the mobile web we have access to the web's services when we are on the move. This book demystifies the tools that we use when interacting with the web, and gives the reader a detailed overview of where we are and where we are going in terms of search engine and web navigation technologies.

Buy Now

Click here for more information about An Introduction to Search Engines and Web Navigation

Read More...

Pro Hadoop (Expert's Voice in Open Source) Review

Pro Hadoop (Expert's Voice in Open Source)
Average Reviews:

(More customer reviews)
The reason why I say this book's still a Good Buy is because Jason Venner has used Hadoop in several scenarios, and this book contains a lot of practical and time-saving tips on what mistakes to avoid or how to troubleshoot problems, making it an especially good book for Hadoop newbies. His materials on Testing and Debugging MapReduce Applications are also a value-add.
Chapter One provides detailed instructions on how to install Hadoop and how to run a test to verify that everything went fine. The author mentions that Hadoop 0.19 works best with Sun's JDK 1.6 and that although Hadoop will work on Windows with Cygwin installed, you have to be careful when specifying file paths.
Chapters Two and Three introduce basic concepts pertaining to MapReduce Jobs and Multimachine Clusters, respectively, and how "master" and "slave" nodes are configured. Chapter Four teaches you how to install, configure, and troubleshoot Hadoop Distributed File System.
Chapters Five and Six provide tutorials on the different types of inputs and outputs that a Hadoop MapReduce job can handle, and how to tune MapReduce jobs.
Chapter Seven is an excellent tutorial on how to unit test and debug MapReduce jobs, while Chapter Eight discusses more advanced MapReduce techniques for addressing more complex application requirements.
Chapter Nine walks you through the evolution of a (somewhat boring) real-world application, discussing rationales behind design changes, etc. Chapter 10 provides a few descriptive paragraphs each for various projects related to Hadoop (e.g., Pig, HBase, Mahout, ZooKeeper,etc). Finally, Appendix A is a detailed discussion of the JobConf API, JobConf being the object that controls information relating to a MapReduce job.

Click Here to see more reviews about: Pro Hadoop (Expert's Voice in Open Source)

It's a very safe bet that cloud computing interest increases and in fact, a near certainty. A recent estimate from Merrill Lynch suggested a $95 billion annual market by 2013. Hadoop is at the center of cloud computing: it is one of the most searched-for, documented, and prevalent form of cloud computing data access, and even in its pre-final form is already rich enough to support a consulting business model. Anyone wishing to investigate an enterprise-level cloud computing solution will need a Hadoop book to at least investigate the possibilities.Hadoop is one of the tools that's driving today's developers to build tomorrow's Software as a Service (SaaS)-based and driven Internet applications, invested in by Yahoo, Microsoft, Google, Amazon and more. With Hadoop, developers can build tomorrow's data centers, or the next 'Microsoft Office" or 'Google Apps" - applications that are fully hosted on the Web, not the desktop - and much more.Pro Hadoop will be the first to market with a professional guide and reference to getting up to speed with using, developing and working with Hadoop, an open source Java-based cloud computing framework and platform backed by Yahoo. This book will likely time with Hadoop 1.0 release in June 2009, also around JavaOne, the world's largest Java conference at around 15,000 attendees on average.

Buy Now

Click here for more information about Pro Hadoop (Expert's Voice in Open Source)

Read More...

Internet Searches for Vetting, Investigations, and Open-Source Intelligence Review

Internet Searches for Vetting, Investigations, and Open-Source Intelligence
Average Reviews:

(More customer reviews)
I have been a private investigator for more than 25 years. When I started in this business computers were scarce and the Internet had not yet been commercialized. Everything was done with a pencil and telephone. I have a lot of experience so it is hard for me to find books that are useful. I buy books hoping to learn one or two new things. Edward J. Appel, a retired FBI agent is the author. His company has done work for me so I knew the quality of his work and assumed his book would be at the same level as his investigative efforts. I wasn't disappointed.
The book is 320 pages and broken into 4 sections. Appel begins with a chapter about behavior and technology which orients the investigator/analyst on the growth and use of the Internet. He discusses the usefulness of the Internet as an investigative tool and the transformations the Internet is making through social and technological advances. Along with the benefits for investigators comes the darker side of the web, and Appel examines its criminal exploitation.
The first section provides a great introduction and is geared toward corporate investigators and security personnel tasked with monitoring IT systems, vetting employees and guarding company intellectual property. In Section 2 Appel outlines legal and policy issues related to using information from the Internet in investigations. He identifies liability and privacy issues and laws addressing these concerns such as the Fair Credit Reporting Act. He also dedicates a chapter to litigation, defamation, and invasion of privacy torts.
Appel develops a framework for preparation and planning of successful Internet research. He covers basic information about search engines, metasearch engines, social networking sites and search terms.
I found the chapters Automation of Searching and Internet Intelligence Reporting to be the most enlightening. Reducing time spent searching and analyzing information, seems like a no-brainer. But I would bet the number of investigators who have looked for automated solutions to their collection efforts is small. The Internet Intelligence Reporting section recommends the format and organization of a report as well as what to include and how to cite sources.
Internet Searches for Vetting, Investigations, and Open-Source Intelligence, like many trade publications, is on the pricey side. However it delivers with content and is an easy read. It met my criteria for a successful investigations book because I learned much more than two new things.


Click Here to see more reviews about: Internet Searches for Vetting, Investigations, and Open-Source Intelligence

In the information age, it is critical that we understand the implications and exposure of the activities and data documented on the Internet. Improved efficiencies and the added capabilities of instant communication, high-speed connectivity to browsers, search engines, websites, databases, indexing, searching and analytical applications have made information technology (IT) and the Internet a vital issued for public and private enterprises. The downside is that this increased level of complexity and vulnerability presents a daunting challenge for enterprise and personal security.Internet Searches for Vetting, Investigations, and Open-Source Intelligence provides an understanding of the implications of the activities and data documented by individuals on the Internet. It delineates a much-needed framework for the responsible collection and use of the Internet for intelligence, investigation, vetting, and open-source information. This book makes a compelling case for action as well as reviews relevant laws, regulations, and rulings as they pertain to Internet crimes, misbehaviors, and individuals' privacy. Exploring technologies such as social media and aggregate information services, the author outlines the techniques and skills that can be used to leverage the capabilities of networked systems on the Internet and find critically important data to complete an up-to-date picture of people, employees, entities, and their activities. Outlining appropriate adoption of legal, policy, and procedural principles-and emphasizing the careful and appropriate use of Internet searching within the law-the book includes coverage of cases, privacy issues, and solutions for common problems encountered in Internet searching practice and information usage, from internal and external threats. The book is a valuable resource on how to utilize open-source, online sources to gather important information and screen and vet employees, prospective employees, corporate partners, and vendors.

Buy Now

Click here for more information about Internet Searches for Vetting, Investigations, and Open-Source Intelligence

Read More...

HTML5 Geolocation Review

HTML5 Geolocation
Average Reviews:

(More customer reviews)
Upshot: Here is your intermediate guide to HTML5 Geolocation-centric APIs. Pros: Another straightforward O'Reilly intro to HTML5-like cutting edge hotness. Cons: None really. Code might change over time, so get the eBook for free updates.

If you need a quick introduction to Geolocation APIs available from Google, as well as ArcGIS, this is it. However this book is not for beginners. You should be very comfortable coding HTML or Javascript as the as there is lengthy code (not just snippets) to get this magic to happen. It is nice that the code that does the magic has been thoughtfully called out and explained by the author.
The author also does a nice job of outlining just *what* Geolocation is (it's not just a flat 2d map), what resources are available, as well as what resources can/need to be saved, and what you can do with that information. The breakdowns of the geo-specific code are very straightforward, and there is a LOT of code, so I recommend getting the eBook version.
Finally there's a section on marketing & privacy and whether that still matters to the younger generation --and how/why all this social network-spacial relationship stuff works in context.
This book is NOT for absolute beginners. But if you're already comfortable with HTML/Javascript, and SOME programming concepts, this book is a nice complement to HTML5 Up & Running by Mark Pilgrim.
Disclosure: I received the eBook download from O'Reilly for review purposes.

Click Here to see more reviews about: HTML5 Geolocation


Truly revolutionary: now you can write geolocation applications directly in the browser, rather than develop native apps for particular devices. This concise book demonstrates the W3C Geolocation API in action, with code and examples to help you build HTML5 apps using the "write once, deploy everywhere" model. Along the way, you get a crash course in geolocation, browser support, and ways to integrate the API with common geo tools like Google Maps.

Learn how geo information is gathered from different sources, depending on the device
Discover how coordinate systems work, including geodetic systems and datums
Use the API to collect location information from a user's browser with JavaScript code
Place geo information on a map using the Google Maps or ArcGIS JavaScript APIs
Save geo data with databases, the Keyhole Markup Language, or the shapefile format
Be familiar with several practical uses for geo data, such as geomarketing, geosocial, geotagging, and geo-applications


Buy Now

Click here for more information about HTML5 Geolocation

Read More...

Mining the Social Web: Analyzing Data from Facebook, Twitter, LinkedIn, and Other Social Media Sites Review

Mining the Social Web: Analyzing Data from Facebook, Twitter, LinkedIn, and Other Social Media Sites
Average Reviews:

(More customer reviews)
Mining the Social Web does a great job of introducing a wide variety of techniques and wealth of resources for exploring freely available social data and personal information. If you are willing to spend the time tinkering with the examples, the book is pure fun. It offers a nice compliment to Segaran's Programming Collective Intelligence: Building Smart Web 2.0 Applications. The two books overlap but where they do offer different perspectives and explanations of common techniques (e.g., TF-IDF, cosine similarity, Jaccard index). If you are well-versed in data mining the web you may find much of the discussion familiar. If you have only been casually engaged to date, your toolbox will fill quickly.
In order to work with the book's examples related to LinkedIn and Facebook you really need to have a robust collection of connections. In terms of the source code itself, most of it worked as is. I wasn't able to install the Buzz library which limited my interaction with material in chapter 7 and opted to not get involved with the LinkedIn or Facebook but found the discussions around them easy to follow. By far my favorite chapter in the book was chapter 8, "Blogs et al.: Natural Language Processing (and Beyond)..." It was quite fascinating and caused my reading list to grow considerably.

Click Here to see more reviews about: Mining the Social Web: Analyzing Data from Facebook, Twitter, LinkedIn, and Other Social Media Sites


Facebook, Twitter, and LinkedIn generate a tremendous amount of valuable social data, but how can you find out who's making connections with social media, what they're talking about, or where they're located? This concise and practical book shows you how to answer these questions and more. You'll learn how to combine social web data, analysis techniques, and visualization to help you find what you've been looking for in the social haystack, as well as useful information you didn't know existed.

Each standalone chapter introduces techniques for mining data in different areas of the social Web, including blogs and email. All you need to get started is a programming background and a willingness to learn basic Python tools.

Get a straightforward synopsis of the social web landscape
Use adaptable scripts on GitHub to harvest data from social network APIs such as Twitter, Facebook, and LinkedIn
Learn how to employ easy-to-use Python tools to slice and dice the data you collect
Explore social connections in microformats with the XHTML Friends Network
Apply advanced mining techniques such as TF-IDF, cosine similarity, collocation analysis, document summarization, and clique detection
Build interactive visualizations with web technologies based upon HTML5 and JavaScript toolkits

"Let Matthew Russell serve as your guide to working with social data sets old (email, blogs) and new (Twitter, LinkedIn, Facebook). Mining the Social Web is a natural successor to Programming Collective Intelligence: a practical, hands-on approach to hacking on data from the social Web with Python." --Jeff Hammerbacher, Chief Scientist, Cloudera

"A rich, compact, useful, practical introduction to a galaxy of tools, techniques, and theories for exploring structured and unstructured data." --Alex Martelli, Senior Staff Engineer, Google


Buy Now

Click here for more information about Mining the Social Web: Analyzing Data from Facebook, Twitter, LinkedIn, and Other Social Media Sites

Read More...

Google Hacks: Tips & Tools for Finding and Using the World's Information Review

Google Hacks: Tips and Tools for Finding and Using the World's Information
Average Reviews:

(More customer reviews)
Few of today's web-savvy would contest Google's superiority among search engines. Behind the austere and simple interface lies a wealth of information just waiting to be tapped. Until now however, tapping all that information and power would likely require scanning dozens of websites hunting down tips for making the most out of Google. Fortunately, Tara Calishain, Rael Dornfest and their colleagues have done most of the legwork for us in O'Reilly's Google Hacks.
Google Hacks is another in O'Reilly's Hacks series, "Industrial Strength Tips and Tools". In this case, 100 recipes for just about every imaginable use for Google. O'Reilly uses the term 'hack' in a positive way, meaning a clever technical feat or trick, as opposed to the negative connotation associated with those blackhats who break into computer systems for fun and for profit. Each "hack" is a stand-alone recipe demonstrating some aspect of using Google to find just what you're looking for. Most hacks also contain cross-references to other relevant hacks in the book, so you really don't have to read it from cover to cover. You could start with whatever interests you, and go from there.
The book is divided into several chapters, each of which contains several hacks. The first few chapters are targeted at the general end-user, describing in detail all of the various syntaxes you can use when searching with Google, as well as introducing the various topical collections (U.S. government, Linux, Mac, etc.), and other tools (Google Groups, Google News, etc.,) available. The authors are careful to point out where the various syntax pieces are incompatible, and which syntax features are available with which services. Also covered are various tools you can use to (legally) 'scrape' Google search results for further analysis. These chapters will be useful for just about anyone who uses Google. Some of the material (such as directly manipulating URLS to tweak search results and custom HTML forms) may be beyond the reach of some newbies. A general understanding of URLs, HTML and CGI scripting will be helpful in making use of most of the book.
The next few chapters are targeted more to developers and propeller-heads, describing the Google Web Service API, as well as providing dozens of scripts (mostly in Perl) for manipulating Google's index via its XML interface. Newbies and the casual user might find all this a bit overwhelming, but anyone with a Perl interpreter could potentially use these scripts to their advantage. One chapter also provides examples of using the API in various other languages including PHP, Java, Python, C#/.NET, and VB.NET. There are enough examples here of using the API in various fashions to get anyone with a sense of programming plenty of starting off points for whatever project they may imagine with Google's wealth of information.
The next to last chapter involves a handful of pranks, games, other oddities you can do with Google. Fool your friends with 0-result searches, let Google write poetry or a recipe for you. Draw pictures with Google Groups, or see just how good you are at Google-Whacking. This is the chapter for all of you who have way too much time on your hands ;-).
The last chapter in the book is targeted towards webmasters and offers several tips not only on getting your website well-placed in Google's search rankings, but also general help on getting traffic to your site in the first place. The authors also discuss strategies for using Google's AdWords system to the advantage of your business.
Overall, the book is very readable, and easy to move through (well, for a geek anyways). Each hack is self-contained, and can be read in a few minutes. Read it near your computer, as you'll likely be wanting to try some of these hacks out as you read them. As for its usefulness, I'm already using things I learned in the book on a regular basis to my daily advantage. However, if you're not more than a casual user of Google, all the scripts and API-speak might be overkill for your needs. The first few and last chapters probably justify the Amazon price for most users, however.
The book isn't perfect, though. I did find a few typographical errors scattered through the text, but they weren't prevalent enough to be too distracting. Also, with coverage of such a moving target as a major Internet property like Google, there will likely be links and even certain hacks that may not work, and features that change with time. Finally, the idea of narrowing down your search results to a manageable number surfaces often. In my opinion, what's important is not so much how many search results are found, but rather, whether or not Google can get me what I'm looking for within the first page or two of results, which it usually does, and which is why I use Google in the first place. The real value of the book shows itself on those occasions where Google doesn't necessarily get you where you want to be on the first shot.
In summary, true to its cover graphic, Google Hacks will provide you with a large number of tools to get the most out of Google, whether for serious research, casual browsing, procrastination activities, or just plain old fun.

Click Here to see more reviews about: Google Hacks: Tips & Tools for Finding and Using the World's Information



Buy Now

Click here for more information about Google Hacks: Tips & Tools for Finding and Using the World's Information

Read More...

Social Network Data Analytics Review

Social Network Data Analytics
Average Reviews:

(More customer reviews)
This is a very interesting book for both researchers and practitioners in computer science who work in the area of data mining and want to learn the state-of-the-art in social network data analytics. The book provides good coverage of the subject area by focusing on popular research topics, such as the study of the statistical properties that are apparent in "typical" social networks, the problems of community detection and social influence analysis, the expert-location discovery problem, the privacy issues that arise in the context of social networks, as well as visualization techniques, text mining techniques and social tagging. The emerging area of integrating sensors and social networks is also examined. Each chapter of the book contains numerous bibliographic references that will guide readers who are interested in particular topics to explore these topics in more depth. Overall, I highly recommend this book!

Click Here to see more reviews about: Social Network Data Analytics


Social network analysis applications have experienced tremendous advances within the last few years due in part to increasing trends towards users interacting with each other on the internet. Social networks are organized as graphs, and the data on social networks takes on the form of massive streams, which are mined for a variety of purposes.

Social Network Data Analytics covers an important niche in the social network analytics field. This edited volume, contributed by prominent researchers in this field, presents a wide selection of topics on social network data mining such as Structural Properties of Social Networks, Algorithms for Structural Discovery of Social Networks and Content Analysis in Social Networks. This book is also unique in focussing on the data analytical aspects of social networks in the internet scenario, rather than the traditional sociology-driven emphasis prevalent in the existing books, which do not focus on the unique data-intensive characteristics of online social networks. Emphasis is placed on simplifying the content so that students and practitioners benefit from this book.

This book targets advanced level students and researchers concentrating on computer science as a secondary text or reference book. Data mining, database, information security, electronic commerce and machine learning professionals will find this book a valuable asset, as well as primary associations such as ACM, IEEE and Management Science.


Buy Now

Click here for more information about Social Network Data Analytics

Read More...

Web Dragons: Inside the Myths of Search Engine Technology (The Morgan Kaufmann Series in Multimedia Information and Systems) Review

Web Dragons: Inside the Myths of Search Engine Technology (The Morgan Kaufmann Series in Multimedia Information and Systems)
Average Reviews:

(More customer reviews)
Anyone who has used a search engine - and who hasn't? - should read this book. It's a very approachable and coherent look at how search engines work, and their role in our information society. The theme of the books is sobering: for many people, access to the internet is through one site - a search engine that has become a "dragon" guarding access to a mine of information.
I particularly like the writing style; the (somewhat dry) humour and intriguing stories are engaging, and on-line tools that we use daily are shown in a new light. The book is suitable for the lay person, but is still engaging for the technically inclined. It provides details about how search engines really work using meaningful examples and illustrations, as well as the exposing the social implications.
Some of the important issues covered include the borderline between spam and content-targeted advertising, determining the authority of web pages compared with their popularity, and issues such as censorship, privacy and access to information. Topics range from the great library of Alexander to the most common misspellings of "Britney Spears" typed into Google.
This book looks set to become part of the computing canon, and would sit equally well on a shelf of technical books or a coffee table. You won't be able to use it to implement your next search engine - it doesn't go into that level of detail. But it's a thought-provoking read, and would be a great gift for the curious or technically inclined people in your life.

Click Here to see more reviews about: Web Dragons: Inside the Myths of Search Engine Technology (The Morgan Kaufmann Series in Multimedia Information and Systems)



Buy Now

Click here for more information about Web Dragons: Inside the Myths of Search Engine Technology (The Morgan Kaufmann Series in Multimedia Information and Systems)

Read More...

Hibernate Search in Action Review

Hibernate Search in Action
Average Reviews:

(More customer reviews)
I just finished reading Hibernate Search in Action, and I loved it. I should point out that I was the porter of Hibernate Search to NHibernate Search, so I had some previous expertise in the topic. In addition to that, I approached this book at an angle completely orthogonal to the expected audience. Unlike most "in Action" books, I did not intend to make immediate use of the code and approaches suggested in the book. Instead, I looked to the book as a way to deepen my understanding of the tool and how it works.
I am impressed, massively so, that it did so well in this regard for someone who has gone through the entire source code of the project several times.
I'll not bore you with the actual details, you can get the actual content summary of off the site. From my perspective, after reading this book I know that I am going to take a completely different approach for most complex search scenarios, and I think that I have the practical theoretical knowledge to deal with it.
I highly recommend the book if you actually need to deal with Hibernate Search, but I would recommend it to people who are not using it, because it contains some important eye opening concepts if you are not used to full text search tools capabilities. As a nice bonus, I was able to take the information in the book and use it to discuss a problem the customer was having, ending up with something that I consider far superior of the solution that they currently employ.

Click Here to see more reviews about: Hibernate Search in Action


Enterprise and web applications require full-featured, "Google-quality" search capabilities, but such features are notoriously difficult to implement and maintain. Hibernate Search builds on the Lucene feature set and offers an easyto- implement interface that integrates seamlessly with Hibernate-the leading data persistence solution for Java applications.

Hibernate Search in Action introduces both the principles of enterprise search and the implementation details a Java developer will need to use Hibernate Search effectively. This book blends the insights of the Hibernate Search lead developer with the practical techniques required to index and manipulate data, assemble and execute search queries, and create smart filters for better search results. Along the way, the reader masters performance-boosting concepts like using Hibernate Search in a clustered environment and integrating with the features already in your applications.

This book assumes you're a competent Java developer with some experience using Hibernate and Lucene.


Buy Now

Click here for more information about Hibernate Search in Action

Read More...

Algorithms of the Intelligent Web Review

Algorithms of the Intelligent Web
Average Reviews:

(More customer reviews)
I have always had an interest in AI, machine learning, and data mining but I found the introductory books too mathematical and focused mostly on solving academic problems rather than real-world industrial problems. So, I was curious to see what this book was about.

I have read the book front-to-back (twice!) before I write this report. I started reading the electronic version a couple of months ago and read the paper print again over the weekend. This is the best practical book in machine learning that you can buy today -- period. All the examples are written in Java and all algorithms are explained in plain English. The writing style is superb! The book was written by one author (Marmanis) while the other one (Babenko) contributed in the source code, so there are no gaps in the narrative; it is engaging, pleasant, and fluent. The author leads the reader from the very introductory concepts to some fairly advanced topics. Some of the topics are covered in the book and some are left as an exercise at the end of each chapter (there is a "To Do" section, which was a wonderful idea!). I did not like some of the figures (they were probably made by the authors not an artist) but this was only a minor aesthetic inconvenience.
The book covers four cornerstones of machine learning and intelligence, i.e. intelligent search, recommendations, clustering, and classification. It also covers a subject that today you can find only in the academic literature, i.e. combination techniques. Combination techniques are very powerful and although the author presents the techniques in the context of classifiers, it is clear that the same can be done for ecommendations -- as the Bell Korr team did for the Netflix prize.
I work in a financial company and a number of people that I work with have PhD degrees in mathematics and computer science. I found the book so fascinating that I asked them to have a look. They had nothing but praise for this book. The consensus is that everything is explained in the simplest possible way, with clarity but without sacrificing accuracy. As one of them told me, this is a major step forward in teaching AI techniques and introducing the field to millions of developers around the world. Even for experts in the field and experienced software engineers, there are important insights in almost every chapter.
We had tried to write a software library, for a small project, that analyzes log files and assesses IT risk (e.g. probability of intrusion; preemptive alerts on application performance issues, and so on) based on Segaran's book "Programming collective intelligence". We spend about six weeks trying to find how to match what was in Segaran's book and what we wanted to do but we did not find the depth and clarity that was required. On top of that, Segaran used Python so the code had to be rewritten and things didn't quite work as expected! We are now using the code from Marmanis' book and our code analyzes apache and weblogic log files in order to assess risk! It just works! We wrote the code in one week! We would not have been able to succeed without reading this book.
Clearly, I am deeply impressed. This is an outstanding book; it was not just useful, it was inspiring! It is a "must have" book for every Java developer.
The content of the book includes:
* the PageRank algorithm; a content based algorithm similar to PageRank to which the author coined the term "DocRank" because it applies to Word, PDF, and other documents rather than Web pages; search improvements based on probabilistic methods (Naive Bayes); precision, recall, F1-score, and ROC curves;
* collaborative filtering as well as content based recommendations;
* k-means, ROCK, DBSCAN for clustering; the best explanation about the "curse of dimensionality" ever! I finally learned what this mystic term means!
* Bayesian classification; declarative programming (through the Drools rules engine); introduction to neural networks; decision trees
* Comparing and Combining classifiers: McNemar's test; Cochran'sQ test; F-test; Bagging; Boosting; general classifier ensembles
Buy it, read it, enjoy it, and use it!


Click Here to see more reviews about: Algorithms of the Intelligent Web


Web 2.0 applications provide a rich user experience, but the parts you can't see are just as important-and impressive. They use powerful techniques to process information intelligently and offer features based on patterns and relationships in data. Algorithms of the Intelligent Web shows readers how to use the same techniques employed by household names like Google Ad Sense, Netflix, and Amazon to transform raw data into actionable information.

Algorithms of the Intelligent Web is an example-driven blueprint for creating applications that collect, analyze, and act on the massive quantities of data users leave in their wake as they use the web. Readers learn to build Netflix-style recommendation engines, and how to apply the same techniques to social-networking sites. See how click-trace analysis can result in smarter ad rotations. All the examples are designed both to be reused and to illustrate a general technique- an algorithm-that applies to a broad range of scenarios.

As they work through the book's many examples, readers learn about recommendation systems, search and ranking, automatic grouping of similar objects, classification of objects, forecasting models, and autonomous agents. They also become familiar with a large number of open-source libraries and SDKs, and freely available APIs from the hottest sites on the internet, such as Facebook, Google, eBay, and Yahoo.


Buy Now

Click here for more information about Algorithms of the Intelligent Web

Read More...