Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

It strikes me that both in the article and in most comments people have no idea of what they are talking about, and yet they boldly carry on.

"The index"? Feature extraction is the most complex part of almost any machine learning algorithm, and search is no different. Indexing full text documents is a really difficult task, especially if you take inflected languages into account (English is particularly easy).

I don't see a way to "open the index" without disclosing and publishing a huge amount of highly complex code, that also makes use of either large dictionaries, or huge amounts of statistical information. It's not like you can just write a quick spec of "the index" and put it up on github.

FWIW, I run a startup that wrote a search engine for e-commerce (search as a service).



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: